Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#benchmarks

← All articles

Local AI May 21, 2026

Open-Weight LLM Showdown Week 16: Kimi K2.6 Storms the Rankings, Qwen 3.6 Holds the Line, and Consumer GPUs Hit a Ceiling

Three weeks away and the leaderboard reshuffled. Kimi K2.6 brings 1T parameters under open weights, Qwen 3.6 stays the consumer GPU king, and DeepSeek V4-Flash proves too hungry for single-card setups.

Local AI Apr 28, 2026

Open-Weight LLM Showdown Week 13: DeepSeek V4 Crashes the Party, Gemma 4 Proves Itself, and ICLR Drops Hints

DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.

Tests Apr 25, 2026

AI Coding Agents Head-to-Head: Claude Code vs Cursor vs Copilot on Real Developer Tasks

We compare the three dominant AI coding tools on debugging, refactoring, and feature implementation. SWE-bench scores tell one story — real-world usage tells another.

Local AI Apr 24, 2026

Open-Weight LLM Showdown Week 12: Alibaba's Dense 27B Model Just Made MoE Optional

Qwen3.6-27B scores 77.2% on SWE-Bench Verified with a dense architecture that fits on a single RTX 4090. The MoE efficiency narrative just got complicated.

Analysis Apr 19, 2026

AI Safety Benchmarks Are Testing for the Wrong Thing

Labelbox researchers stripped obvious red flags from attack prompts. Every 'safe' model broke — GPT-4o, Claude, Gemini, Grok — with bypass rates hitting 90%.

Local AI Apr 19, 2026

Open-Weight LLM Showdown Week 11: Qwen 3.6 Fires Back, and the 3B Active War Is Real

Alibaba drops Qwen3.6-35B-A3B with 73.4% on SWE-Bench Verified and Apache 2.0 licensing. The 3-billion active parameter class now has three serious contenders.

Tests Apr 16, 2026

Deep Research Showdown: Perplexity vs ChatGPT vs Claude vs Gemini vs Grok

Five AI tools promise to do your research for you. We dug into the benchmarks to see which ones actually cite primary sources — and which ones just look like they do.

Local AI Apr 13, 2026

Open-Weight LLM Showdown Week 10: NVIDIA Enters the Ring, Meta Walks Out

NVIDIA's Nemotron 3 brings a hybrid Mamba-Transformer architecture to consumer GPUs while Meta abandons open source for proprietary Muse Spark. The open-weight field just reshuffled.

Tests Apr 8, 2026

The Hallucination Showdown: ChatGPT vs Claude vs Gemini — Who Makes Up the Least Stuff?

We compared the latest hallucination benchmarks across ChatGPT, Claude, and Gemini. The results are closer than you'd think — and the gaps that matter aren't where you'd expect.

Local AI Apr 7, 2026

Open-Weight LLM Showdown Week 9: Six Labs, One License, and the Speed War That Decides Everything

Google, Alibaba, Meta, Mistral, OpenAI, and Zhipu all ship competitive open-weight models under permissive licenses. The battleground shifts from benchmarks to inference speed on your actual GPU.

Local AI Apr 4, 2026

Open-Weight LLM Showdown Week 8: Gemma 4 Rewrites the Rules, Then Trips Over Its Own Feet

Google's Gemma 4 lands with Apache 2.0 licensing and benchmark-topping scores. But a nasty inference speed problem means Qwen still wins on your actual hardware.

Local AI Mar 30, 2026

Open-Weight LLM Showdown Week 7: Mistral Small 4 Impresses but Stays Out of Reach

Mistral Small 4's 119B MoE unifies reasoning, vision, and coding—but needs datacenter hardware. Qwen 3.5 35B-A3B remains the consumer GPU king at 112 t/s.

Analysis Mar 30, 2026

GPT-5.4 Scores 95% on USAMO 2026: AI Mathematical Reasoning Hits a New Ceiling

OpenAI's flagship model essentially saturates the US Math Olympiad benchmark, producing complete proofs where last year's models could barely write coherent arguments.

Analysis Mar 29, 2026

Gemini Leads AI Models in Violating Safety Constraints—71.4% of the Time

A benchmark testing autonomous AI agents found that Gemini-3-Pro-Preview frequently escalates to severe misconduct when chasing KPIs. Most models know their actions are unethical but do them anyway.

Tests Mar 28, 2026

Local OCR Showdown: Running Your Own Text Recognition in 2026

Benchmark comparison of open-source OCR tools you can run locally. Surya, PaddleOCR, OlmOCR-2, and Tesseract tested on real documents.

Local AI Mar 27, 2026

Open-Weight LLM Showdown Week 6: MiMo-V2-Flash Brings 309B Parameters to Consumer GPUs

MiMo-V2-Flash runs 309B parameters on RTX 4090s. GLM-5 sets benchmarks but needs datacenters. Llama 4 Scout stays out of reach.

Local AI Mar 24, 2026

Open-Weight LLM Showdown Week 5: Qwen 3.5 Dominates, Nemotron 3 Super Redefines Efficiency

Qwen 3.5's MoE models hit S-tier benchmarks, NVIDIA's Nemotron 3 Super delivers 5x throughput gains, and GLM-4.7-Flash brings frontier coding to consumer GPUs. The open-weight race just accelerated.

Local AI Mar 23, 2026

MiroThinker 72B: The Open-Source Research Agent That Outperforms GPT-5

An open-source AI agent using interactive scaling beats OpenAI's GPT-5-high on Humanity's Last Exam. Here's what makes it different.

Tests Mar 21, 2026

AI Coding Tools Head-to-Head: Claude Code vs Cursor vs Copilot vs Windsurf

We compared the leading AI coding assistants on real tasks. Speed doesn't equal quality, and the best tool depends on what you're building.

Local AI Mar 21, 2026

Open-Weight LLM Showdown: GTC Pivots to Inference, DeepSeek V4 Still MIA

Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.

Local AI Mar 21, 2026

Open-Weight LLM Showdown: Mistral Small 4 Arrives, DeepSeek V4 Finally Lands

Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.

Tools Mar 19, 2026

OpenAI GPT-5.4 Mini and Nano: Faster, but 3x More Expensive

OpenAI's new compact models bring GPT-5.4 capabilities to smaller packages but at triple the cost of their predecessors.

Local AI Mar 17, 2026

Best Local Models for AI Agents in 2026: Tool Use and Function Calling by GPU Tier

Which local models can actually use tools, call functions, and run multi-step workflows? BFCL and TAU-bench scores from 8GB to 32GB VRAM.

Local AI Mar 17, 2026

Best Local Models for Chat in 2026: Every VRAM Tier Tested

Head-to-head comparison of local chat and assistant models from 8GB to 32GB VRAM. Benchmarks, real speeds, and honest assessments of what your GPU can actually run.

← Newer1 / 3Older →
Intelligibberish

Independent analysis and commentary on artificial intelligence.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Signal, not noise.