Open-Weight LLM Showdown: Mistral Small 4 Arrives, DeepSeek V4 Finally Lands
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
Category
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
GTC 2026's biggest announcements were open-source. Nemotron 3 Super runs locally on RTX PCs, LTX 2.3 generates 4K video with audio, and vLLM hits production grade.
TikTok's parent company just open-sourced a powerful framework for running coordinated AI agents on your own hardware. Here's what it does and how to set it up.
Which local models can actually use tools, call functions, and run multi-step workflows? Function-calling and TAU-bench picks from 8GB to 32GB VRAM.
Head-to-head comparison of local chat and assistant models from 8GB to 32GB VRAM. Current picks: Qwen3.5, Gemma 4, GPT-OSS, Qwen3.6, and GLM-4.7-Flash.
Which open-weight coding model to run locally? HumanEval and SWE-bench picks from 8GB to 32GB GPUs. Qwen2.5-Coder, Qwen3.6, Devstral, KAT-Coder.
Local TTS and STT by VRAM tier: Parakeet, Canary, MOSS-Transcribe-Diarize, Step-Audio-EditX, Fish Audio S2 Pro and Kokoro, and the licence each ships.
TranslateGemma, NLLB-200, Aya Expanse and Qwen3.5 by VRAM tier, with the licence terms that decide whether you can ship what you run.
Local image analysis, OCR, and visual reasoning from 8GB to 32GB VRAM. Qwen3.5 replaces Qwen3-VL at most tiers, and 16GB stays unresolved.
Georgi Gerganov's team is now at Hugging Face, unifying the model hub with the inference engine that powers Ollama, LM Studio, and the entire local AI ecosystem.
Local AI on a 12GB GPU: chat, coding, vision, speech and agents for RTX 3060 12GB or RTX 4070. Current picks, per-quant weight sizes, honest limits.
Local AI on a 16GB GPU: chat, coding, translation, speech and agents for RTX 4060 Ti, RTX 5060 or Arc A770, and why the vision tier stays unresolved.
Local AI on a 24GB GPU: chat, coding, vision, speech and agents for RTX 3090 or RTX 4090. Current picks, per-quant weight sizes, and an open runtime bug.
Local AI on a 32GB GPU: chat, coding, vision, speech and agents on an RTX 5090. Current picks, per-quant weight sizes, and what the headroom buys.
Local AI on an 8GB GPU: chat, coding, vision, speech and agents for RTX 4060 or RTX 3070. Current picks, named quantisations, honest limits.
Two weeks after our last roundup, the 5090 benchmarks are in and Qwen 3.5 Small models are running on phones. Here's the real performance picture.
The new MacBook Pro with M5 Max can run large language models entirely on-device, keeping your AI interactions private and offline
Alibaba's new 0.8B to 9B parameter models deliver GPT-class multimodal performance on consumer hardware, with the 9B variant outperforming models 13 times its size
This week's open-source highlights: AI2's hybrid architecture proves transformers need help, autoresearch automates ML experiments overnight, and local inference gets serious upgrades.
Nvidia open-sources a 30B-parameter reasoning model that runs on consumer GPUs with a million-token context window. Here's what makes it different.
AI2 and Lambda trained a hybrid transformer-RNN model that's twice as data-efficient as pure transformers. But can you actually run it locally?
Chinese AI startup MiniMax has released M2.5, an open-weights model matching Claude Opus performance for coding and agentic tasks while costing 95% less to run
GLM-5, Qwen 3.5, DeepSeek V3.2, and MiniMax M2.5 are rewriting the rules. Here's what they actually deliver on consumer hardware.
Ollama's new OpenClaw integration lets you run AI agents locally through WhatsApp, Telegram, or Slack. Here's how it works, what you need, and the security risks nobody mentions.