Open-Weight LLM Showdown: RTX 5090 Finally Delivers, But You Can't Buy One
Two weeks after our last roundup, the 5090 benchmarks are in and Qwen 3.5 Small models are running on phones. Here's the real performance picture.
Tag
Two weeks after our last roundup, the 5090 benchmarks are in and Qwen 3.5 Small models are running on phones. Here's the real performance picture.
xAI's new multi-agent architecture pits four specialized AI agents against each other in real-time debate, claiming 65% fewer hallucinations
New research shows AI can link your Reddit burner account to your real identity for under $4. The era of casual online pseudonymity may be ending.
New research from ETH Zurich and Anthropic shows AI can identify pseudonymous online accounts with 67% accuracy at just $1-4 per person - and there's no easy fix.
Chinese AI startup MiniMax has released M2.5, an open-weights model matching Claude Opus performance for coding and agentic tasks while costing 95% less to run
GLM-5, Qwen 3.5, DeepSeek V3.2, and MiniMax M2.5 are rewriting the rules. Here's what they actually deliver on consumer hardware.
Researchers have created an AI system that can generate text mimicking specific personality traits and mental health conditions. The implications for manipulation and misinformation are troubling.
Stanford and Princeton researchers found Chinese AI models refuse politically sensitive questions at rates up to 60% compared to under 3% for Western models - and the censorship goes beyond training data.
An LLM trained on yeast genetics outperforms commercial tools at optimizing codon sequences for protein production. Five out of six test cases beat existing solutions.
New research from Google and UVA reveals that longer AI reasoning traces actually correlate with wrong answers. The fix: measure how deeply the model thinks, not how much it writes.
Half of the tested AI tools produced prediction models that matched or beat human researchers. A master's student and high schooler built working code in minutes.
Kaspersky research reveals that passwords from ChatGPT, DeepSeek, and Llama lack true randomness. The same prediction capability that makes LLMs useful makes them terrible at generating secure passwords.
Chinese researchers built an AI system using 40+ specialized tools that correctly identifies rare diseases in first attempt 64% of the time vs 55% for experienced physicians.
Zhipu AI releases GLM-5 under MIT license, a frontier model rivaling Claude and GPT-5 while proving China can build top-tier AI without NVIDIA hardware.
Mount Sinai researchers tested 20 LLMs with over a million prompts and found they readily accept false medical claims embedded in clinical-looking documents.
Google's latest model scores 77% on ARC-AGI-2, more than double its predecessor. At $2 per million tokens, it undercuts competitors while outperforming them on most tests.
Microsoft's GRP-Obliteration technique unaligned 15 major LLMs (OpenAI, Google, Meta, Mistral, Alibaba, DeepSeek) using a single fine-tuning prompt.
An Oxford study found AI chatbots diagnose conditions correctly 94.9% of the time on paper, but only 34.5% when talking to actual people. The implications for AI benchmarks extend far beyond medicine.