MLCommons: AI Safety Benchmarking Is Fundamentally Broken
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Tag
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Google's Gemini 3.1 Pro scores 77.1% on ARC-AGI-2. API pricing starts at $2 per million input tokens for prompts up to 200,000 tokens.
A tier-by-tier comparison of the top open-weight LLMs you can run locally, from 8GB laptops to 24GB gaming GPUs to Apple Silicon Macs.
Modern sub-10B models now rival last year's frontier AI on reasoning, tool use, and code. The benchmarks prove it.
Google's upgraded reasoning model finds flaws in peer-reviewed papers, optimizes semiconductor fabrication, and outperforms every frontier model on scientific benchmarks.
An Oxford study found AI chatbots diagnose conditions correctly 94.9% of the time on paper, but only 34.5% when talking to actual people. The implications for AI benchmarks extend far beyond medicine.
A model ID surfaced in Google Cloud this weekend. The AI rumor mill did the rest. We separate the signal from the noise.