Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#benchmarks

← All articles

Analysis Feb 21, 2026

MLCommons: AI Safety Benchmarking Is Fundamentally Broken

Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators

Analysis Feb 20, 2026

Gemini 3.1 Pro More Than Doubles Its ARC-AGI-2 Score

Google's Gemini 3.1 Pro scores 77.1% on ARC-AGI-2. API pricing starts at $2 per million input tokens for prompts up to 200,000 tokens.

Local AI Feb 20, 2026

Local AI Showdown: Best Open-Weight Models for Your Hardware

A tier-by-tier comparison of the top open-weight LLMs you can run locally, from 8GB laptops to 24GB gaming GPUs to Apple Silicon Macs.

Tests Feb 20, 2026

Small Models, Big Brain: When 4 Billion Parameters Match GPT-4

Modern sub-10B models now rival last year's frontier AI on reasoning, tool use, and code. The benchmarks prove it.

Analysis Feb 15, 2026

Google's Gemini 3 Deep Think Solves 18 Unsolved Research Problems and Disproves a Decade-Old Conjecture

Google's upgraded reasoning model finds flaws in peer-reviewed papers, optimizes semiconductor fabrication, and outperforms every frontier model on scientific benchmarks.

Analysis Feb 10, 2026

AI Chatbots Ace Medical Exams but Fail Real Patients: The 60-Point Gap That Should Worry Everyone

An Oxford study found AI chatbots diagnose conditions correctly 94.9% of the time on paper, but only 34.5% when talking to actual people. The implications for AI benchmarks extend far beyond medicine.

Analysis Feb 5, 2026

Claude Sonnet 5 'Fennec' Appeared in Vertex AI Logs

A model ID surfaced in Google Cloud this weekend. The AI rumor mill did the rest. We separate the signal from the noise.

← Newer3 / 3Older →
Intelligibberish

Making sense of AI overwhelm. Independent, self-hosted, no trackers.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Making sense of AI overwhelm.