Open-Weight LLM Showdown: GTC Pivots to Inference, DeepSeek V4 Still MIA
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Tag
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
OpenAI's new compact models bring GPT-5.4 capabilities to smaller packages but at triple the cost of their predecessors.
Which local models can actually use tools, call functions, and run multi-step workflows? Function-calling and TAU-bench picks from 8GB to 32GB VRAM.
Which open-weight coding model to run locally? HumanEval and SWE-bench picks from 8GB to 32GB GPUs. Qwen2.5-Coder, Qwen3.6, Devstral, KAT-Coder.
Head-to-head comparison of local chat and assistant models from 8GB to 32GB VRAM. Current picks: Qwen3.5, Gemma 4, GPT-OSS, Qwen3.6, and GLM-4.7-Flash.
Local TTS and STT by VRAM tier: Parakeet, Canary, MOSS-Transcribe-Diarize, Step-Audio-EditX, Fish Audio S2 Pro and Kokoro, and the licence each ships.
TranslateGemma, NLLB-200, Aya Expanse and Qwen3.5 by VRAM tier, with the licence terms that decide whether you can ship what you run.
Local image analysis, OCR, and visual reasoning from 8GB to 32GB VRAM. Qwen3.5 replaces Qwen3-VL at most tiers, and 16GB stays unresolved.
Local AI on a 12GB GPU: chat, coding, vision, speech and agents for RTX 3060 12GB or RTX 4070. Current picks, per-quant weight sizes, honest limits.
Local AI on a 16GB GPU: chat, coding, translation, speech and agents for RTX 4060 Ti, RTX 5060 or Arc A770, and why the vision tier stays unresolved.
Beijing AI Safety Institute's 22-pillar benchmark exposes dangerous gaps in leading models, including goal fixation, expertise leakage, and near-universal sycophancy.
Local AI on an 8GB GPU: chat, coding, vision, speech and agents for RTX 4060 or RTX 3070. Current picks, named quantisations, honest limits.
Local AI on a 24GB GPU: chat, coding, vision, speech and agents for RTX 3090 or RTX 4090. Current picks, per-quant weight sizes, and an open runtime bug.
Local AI on a 32GB GPU: chat, coding, vision, speech and agents on an RTX 5090. Current picks, per-quant weight sizes, and what the headroom buys.
Two weeks after our last roundup, the 5090 benchmarks are in and Qwen 3.5 Small models are running on phones. Here's the real performance picture.
University of Missouri releases PSBench, a massive benchmark database to help researchers know when AI protein predictions can be trusted
GLM-5, Qwen 3.5, DeepSeek V3.2, and MiniMax M2.5 are rewriting the rules. Here's what they actually deliver on consumer hardware.
Real benchmark results from building a task management dashboard with four leading AI coding tools. Who wins on speed, code quality, and security?
Forget 700B parameter flagships you can't run. Here are the open-weight models that deliver real performance on consumer hardware - with actual benchmarks.
Google and UVA research shows longer AI reasoning traces correlate with wrong answers. The fix: measure how deeply the model thinks, not how much it writes.
Zhipu AI releases GLM-5 under MIT license, a frontier model rivaling Claude and GPT-5 while proving China can build top-tier AI without NVIDIA hardware.
Forget the marketing - here's how the latest open-weight models actually perform on your GPU, from 8GB budget cards to 24GB workstations.
Real benchmark data, developer reviews, and practical tests reveal when each tool wins - and why smart teams use both