How to Choose an Open-Weight Model Family (September 2026)
Qwen, Llama, Mistral, Gemma, DeepSeek, Phi - which family to commit to, what each is good at, and the licence traps that send you back to negotiate.
Tag
Qwen, Llama, Mistral, Gemma, DeepSeek, Phi - which family to commit to, what each is good at, and the licence traps that send you back to negotiate.
Reasoning-capable open-weight models sized by VRAM. DeepSeek-R1, Qwen3 with thinking, QwQ, Phi-4 reasoning, gpt-oss, and the licence traps.
Sixteen downloadable-weight models ranked by capability and licence, with every parameter count and licence re-checked against the Hugging Face API.
UK AISI's July 17 evaluation finds GLM-5.2 and DeepSeek V4-Pro match closed frontier models from 4-7 months ago at a fraction of the cost.
DeepSeek V4, Cohere Command A+, ZAYA1-8B, and NVIDIA Nemotron 3 mark the busiest month for open-weight AI ever.
Three weeks away and the leaderboard reshuffled. Kimi K2.6 brings 1T parameters under open weights, Qwen 3.6 stays the consumer GPU king, and DeepSeek V4-Flash proves too hungry for single-card setups.
DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.
DeepSeek V4 Pro approaches frontier-level performance. Google, Mistral, and Alibaba ship under Apache 2.0. Ollama hits 52 million monthly downloads.
DeepSeek V4 matches Claude Opus on coding at 7x lower cost under MIT license. NVIDIA's Nemotron 3 brings hybrid Mamba-Transformer MoE to the open. Google's TurboQuant cuts KV cache memory by 6x with no retraining.
Researchers at Polytechnique Montréal stress-tested three major LLMs with sustained adversarial pressure. DeepSeek-v3 showed the steepest ethical degradation. None fully recovered.
GPT-4 class models that cost $60 per million tokens in 2024 now cost $8. DeepSeek's open-source disruption triggered the fastest pricing collapse in tech history.
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
Anthropic accuses DeepSeek, Moonshot, and MiniMax of industrial-scale model theft through 24,000 fake accounts. But the company's own copyright history complicates the moral high ground.
Two anonymous Chinese AI models appeared on OpenRouter with no attribution. Developers are split on whether they're DeepSeek V4 or Zhipu GLM-6 testing in stealth mode.
Five missed release windows, a mysterious V4 Lite appearance, and silence from DeepSeek. What's really happening with China's most anticipated AI model?
China's DeepSeek is releasing V4 - a trillion-parameter multimodal model optimized for domestic chips - while blocking US chipmakers and facing distillation accusations from OpenAI and Anthropic.
Stanford and Princeton researchers found Chinese AI models refuse politically sensitive questions at rates up to 60% compared to under 3% for Western models - and the censorship goes beyond training data.
The Chinese AI lab is withholding its flagship model from US chipmakers while the Trump administration alleges it was trained on banned Blackwell chips.
Android malware using Gemini for real-time evasion. A low-skill attacker using Claude and DeepSeek to compromise 600 networks. NIST launches an emergency standards initiative. Welcome to February 2026.
Kaspersky finds DeepSeek, Llama, and ChatGPT all produce password outputs that fail standard strength tests. Prediction capability makes LLMs bad at randomness.
Researchers discovered that displaying an AI model's reasoning process creates a roadmap for attackers. OpenAI's o1 rejection rate dropped from 98% to under 2%.
A source-by-source audit of eight AI assistants, what they collect, how training defaults differ, and which privacy settings users can change.
Research from ELLIS Alicante shows AI reasoning models can autonomously plan and execute attacks that bypass safety guardrails in nearly all other AI systems.