How to Choose an Open-Weight Model Family (September 2026)
Qwen, Llama, Mistral, Gemma, DeepSeek, Phi - which family to commit to, what each is good at, and the licence traps that send you back to negotiate.
Tag
Qwen, Llama, Mistral, Gemma, DeepSeek, Phi - which family to commit to, what each is good at, and the licence traps that send you back to negotiate.
Why general benchmarks like MMLU and GPQA don't predict your results, the index-rebase trap, and a recipe for picking a local model from your own prompts.
Reasoning-capable open-weight models sized by VRAM. DeepSeek-R1, Qwen3 with thinking, QwQ, Phi-4 reasoning, gpt-oss, and the licence traps.
Apache-2.0, MIT, Qwen variants, Llama community terms: what the Hugging Face API actually returns and the per-repo traps that send you back to negotiate.
Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.