Open-Weight LLM Showdown: Mistral Small 4 Arrives, DeepSeek V4 Finally Lands
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
Tag
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
After 12 years at Meta, the Turing Award winner raised the largest European seed round ever to prove the AI industry got it wrong.
Which open-weight coding model to run locally? HumanEval and SWE-bench picks from 8GB to 32GB GPUs. Qwen2.5-Coder, Qwen3.6, Devstral, KAT-Coder.
Head-to-head comparison of local chat and assistant models from 8GB to 32GB VRAM. Current picks: Qwen3.5, Gemma 4, GPT-OSS, Qwen3.6, and GLM-4.7-Flash.
Two weeks after our last roundup, the 5090 benchmarks are in and Qwen 3.5 Small models are running on phones. Here's the real performance picture.
xAI's new multi-agent architecture pits four specialized AI agents against each other in real-time debate, claiming 65% fewer hallucinations
New research shows AI can link your Reddit burner account to your real identity for under $4. The era of casual online pseudonymity may be ending.
New research from ETH Zurich and Anthropic shows AI can identify pseudonymous online accounts with 67% accuracy at just $1-4 per person - and there's no easy fix.
Chinese AI startup MiniMax has released M2.5, an open-weights model matching Claude Opus performance for coding and agentic tasks while costing 95% less to run
GLM-5, Qwen 3.5, DeepSeek V3.2, and MiniMax M2.5 are rewriting the rules. Here's what they actually deliver on consumer hardware.
Researchers have created an AI system that can generate text mimicking specific personality traits and mental health conditions. The implications for manipulation and misinformation are troubling.
Stanford and Princeton researchers found Chinese AI models refuse politically sensitive questions at rates up to 60% compared to under 3% for Western models - and the censorship goes beyond training data.
An LLM trained on yeast genetics outperforms commercial tools at optimizing codon sequences for protein production. Five out of six test cases beat existing solutions.
Google and UVA research shows longer AI reasoning traces correlate with wrong answers. The fix: measure how deeply the model thinks, not how much it writes.
Half of the tested AI tools produced prediction models that matched or beat human researchers. A master's student and high schooler built working code in minutes.
Kaspersky finds DeepSeek, Llama, and ChatGPT all produce password outputs that fail standard strength tests. Prediction capability makes LLMs bad at randomness.
Shanghai researchers built DeepRare, an AI system using 40+ tools that identifies rare diseases 64% vs 55% for experienced physicians on first attempt.
Zhipu AI releases GLM-5 under MIT license, a frontier model rivaling Claude and GPT-5 while proving China can build top-tier AI without NVIDIA hardware.
Mount Sinai researchers tested 20 LLMs with over a million prompts and found they readily accept false medical claims embedded in clinical-looking documents.
Google's Gemini 3.1 Pro scores 77.1% on ARC-AGI-2. API pricing starts at $2 per million input tokens for prompts up to 200,000 tokens.
Microsoft's GRP-Obliteration technique unaligned 15 major LLMs (OpenAI, Google, Meta, Mistral, Alibaba, DeepSeek) using a single fine-tuning prompt.
An Oxford study found AI chatbots diagnose conditions correctly 94.9% of the time on paper, but only 34.5% when talking to actual people. The implications for AI benchmarks extend far beyond medicine.