Open-Weight LLM Showdown: GTC Pivots to Inference, DeepSeek V4 Still MIA
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Tag
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
The Codex maker acquires uv and Ruff, downloaded 126 million times monthly. Open source community watches nervously.
UC Berkeley and UCSF release an open-source radiology AI that processes 3D scans 150x faster than existing models and beats big tech on diagnostic accuracy.
GTC 2026's biggest announcements were open-source. Nemotron 3 Super runs locally on RTX PCs, LTX 2.3 generates 4K video with audio, and vLLM hits production grade.
TikTok's parent company just open-sourced a powerful framework for running coordinated AI agents on your own hardware. Here's what it does and how to set it up.
Two anonymous Chinese AI models appeared on OpenRouter with no attribution. Developers are split on whether they're DeepSeek V4 or Zhipu GLM-6 testing in stealth mode.
Georgi Gerganov's team is now at Hugging Face, unifying the model hub with the inference engine that powers Ollama, LM Studio, and the entire local AI ecosystem.
Two weeks after our last roundup, the 5090 benchmarks are in and Qwen 3.5 Small models are running on phones. Here's the real performance picture.
Five missed release windows, a mysterious V4 Lite appearance, and silence from DeepSeek. What's really happening with China's most anticipated AI model?
Isomorphic Labs' new drug design AI doubles AlphaFold 3's accuracy. Scientists call it groundbreaking. There's just one problem: it's completely proprietary.
Anthropic's Claude Opus 4.6 discovered 14 high-severity bugs in Firefox including a CVSS 9.8 JIT flaw, demonstrating that AI security research can find logic errors traditional tools overlook.
This week's open-source highlights: AI2's hybrid architecture proves transformers need help, autoresearch automates ML experiments overnight, and local inference gets serious upgrades.
Nvidia open-sources a 30B-parameter reasoning model that runs on consumer GPUs with a million-token context window. Here's what makes it different.
AI2 and Lambda trained a hybrid transformer-RNN model that's twice as data-efficient as pure transformers. But can you actually run it locally?
ARXIV OMEGA on Cisco research showing multi-turn jailbreak attacks succeed 93% of the time against open-weight AI models. Just keep talking.
From ACE-Step challenging Suno to Midjourney V8's imminent launch, the creative AI landscape is splitting between open-source freedom and commercial litigation.
Chinese AI startup MiniMax has released M2.5, an open-weights model matching Claude Opus performance for coding and agentic tasks while costing 95% less to run
China's DeepSeek is releasing V4 - a trillion-parameter multimodal model optimized for domestic chips - while blocking US chipmakers and facing distillation accusations from OpenAI and Anthropic.
GLM-5, Qwen 3.5, DeepSeek V3.2, and MiniMax M2.5 are rewriting the rules. Here's what they actually deliver on consumer hardware.
Veea releases a sub-millisecond security proxy for AI agents under MIT license as new research shows 88% of organizations have experienced agent security incidents.
A veteran Google security engineer built a sandbox system that treats AI agents as fundamentally untrusted - and it could be the model for safe agent deployment.
China's MiniMax releases an MIT-licensed model that rivals Claude Opus 4.6 on coding and agentic tasks. The catch: Anthropic accuses MiniMax of stealing Claude's capabilities to build it.
March 2026's open-source AI highlights: Zhipu's GLM-5 rivals GPT-5, OpenAI finally goes open, and the Linux Foundation creates a home for AI agents.