Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#llm

← All articles

Tests Jul 7, 2026

Dartmouth's AI-Graded Textbook Posted 1.30 SD of Learning Gains

Phosphor's LLM-graded textbook quizzes lifted Dartmouth final-exam scores 0.71 to 1.30 SD - but only where students had to type real answers.

Analysis Jul 2, 2026

Flint, Qwen3, and a Year of LLM Groupthink

A tiny Australian startup is fine-tuning a model to break ChatGPT-style groupthink. The 'Time is a river' problem is the symptom.

Local AI May 21, 2026

Open-Weight LLM Showdown Week 16: Kimi K2.6 Storms the Rankings, Qwen 3.6 Holds the Line, and Consumer GPUs Hit a Ceiling

Three weeks away and the leaderboard reshuffled. Kimi K2.6 brings 1T parameters under open weights, Qwen 3.6 stays the consumer GPU king, and DeepSeek V4-Flash proves too hungry for single-card setups.

Local AI Apr 28, 2026

Open-Weight LLM Showdown Week 13: DeepSeek V4 Crashes the Party, Gemma 4 Proves Itself, and ICLR Drops Hints

DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.

Analysis Apr 24, 2026

ICLR 2026 Opens in Rio Under a Cloud: Reviewer Leaks, AI-Written Reviews, and the Papers That Actually Matter

The biggest AI research conference of the year kicks off with 5,355 accepted papers, two controversies that rattled the field, and findings that should worry anyone deploying LLMs in production.

Local AI Apr 24, 2026

Open-Weight LLM Showdown Week 12: Alibaba's Dense 27B Model Just Made MoE Optional

Qwen3.6-27B scores 77.2% on SWE-Bench Verified with a dense architecture that fits on a single RTX 4090. The MoE efficiency narrative just got complicated.

Local AI Apr 19, 2026

Open-Weight LLM Showdown Week 11: Qwen 3.6 Fires Back, and the 3B Active War Is Real

Alibaba drops Qwen3.6-35B-A3B with 73.4% on SWE-Bench Verified and Apache 2.0 licensing. The 3-billion active parameter class now has three serious contenders.

Local AI Apr 13, 2026

Open-Weight LLM Showdown Week 10: NVIDIA Enters the Ring, Meta Walks Out

NVIDIA's Nemotron 3 brings a hybrid Mamba-Transformer architecture to consumer GPUs while Meta abandons open source for proprietary Muse Spark. The open-weight field just reshuffled.

Analysis Apr 12, 2026

Your AI's Safety Training Can Be Surgically Removed at Runtime

Researchers found the exact neurons responsible for refusing harmful requests — then switched them off. No retraining. No fine-tuning. Just geometry.

Analysis Apr 11, 2026

Your AI Chatbot Would Rather Sell You Something Than Help You

Princeton researchers tested 23 LLMs with advertising conflicts of interest. Most chose company profits over user welfare — and treated rich users better.

Analysis Apr 11, 2026

One Line of Code Jailbreaks 11 AI Models — Including the 'Safe' Ones

Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.

Local AI Apr 7, 2026

Open-Weight LLM Showdown Week 9: Six Labs, One License, and the Speed War That Decides Everything

Google, Alibaba, Meta, Mistral, OpenAI, and Zhipu all ship competitive open-weight models under permissive licenses. The battleground shifts from benchmarks to inference speed on your actual GPU.

Local AI Apr 4, 2026

Open-Weight LLM Showdown Week 8: Gemma 4 Rewrites the Rules, Then Trips Over Its Own Feet

Google's Gemma 4 lands with Apache 2.0 licensing and benchmark-topping scores. But a nasty inference speed problem means Qwen still wins on your actual hardware.

Analysis Mar 31, 2026

AI Agent Cracks the Code on Turning CO2 Into Fuel

Tohoku University team uses Catalysis AI Agent to discover universal design principle for copper catalysts that convert carbon dioxide to useful products.

Local AI Mar 30, 2026

Open-Weight LLM Showdown Week 7: Mistral Small 4 Impresses but Stays Out of Reach

Mistral Small 4's 119B MoE unifies reasoning, vision, and coding—but needs datacenter hardware. Qwen 3.5 35B-A3B remains the consumer GPU king at 112 t/s.

Local AI Mar 27, 2026

Open-Weight LLM Showdown Week 6: MiMo-V2-Flash Brings 309B Parameters to Consumer GPUs

MiMo-V2-Flash runs 309B parameters on RTX 4090s. GLM-5 sets benchmarks but needs datacenters. Llama 4 Scout stays out of reach.

Local AI Mar 24, 2026

Open-Weight LLM Showdown Week 5: Qwen 3.5 Dominates, Nemotron 3 Super Redefines Efficiency

Qwen 3.5's MoE models hit S-tier benchmarks, NVIDIA's Nemotron 3 Super delivers 5x throughput gains, and GLM-4.7-Flash brings frontier coding to consumer GPUs. The open-weight race just accelerated.

Analysis Mar 23, 2026

Your Customer Service Chatbot Is a Security Hole

IEEE S&P research finds 10,000+ websites running vulnerable AI chatbot plugins. Attackers can forge conversations, hijack tools, and extract system prompts.

Local AI Mar 21, 2026

Open-Weight LLM Showdown: GTC Pivots to Inference, DeepSeek V4 Still MIA

Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.

Local AI Mar 21, 2026

Open-Weight LLM Showdown: Mistral Small 4 Arrives, DeepSeek V4 Finally Lands

Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.

Analysis Mar 21, 2026

Yann LeCun's $1 Billion Bet: World Models Will Beat LLMs

After 12 years at Meta, the Turing Award winner raised the largest European seed round ever to prove the AI industry got it wrong.

Local AI Mar 17, 2026

Best Local Models for Chat in 2026: Every VRAM Tier Tested

Head-to-head comparison of local chat and assistant models from 8GB to 32GB VRAM. Benchmarks, real speeds, and honest assessments of what your GPU can actually run.

Local AI Mar 17, 2026

Best Local Models for Coding in 2026: Every VRAM Tier Tested

Which open-weight coding model should you run locally? HumanEval, SWE-bench, and real-world tests from 8GB to 32GB GPUs, with setup instructions for IDE integration.

Local AI Mar 16, 2026

Open-Weight LLM Showdown: RTX 5090 Finally Delivers, But You Can't Buy One

Two weeks after our last roundup, the 5090 benchmarks are in and Qwen 3.5 Small models are running on phones. Here's the real performance picture.

← Newer1 / 2Older →
Intelligibberish

Independent analysis and commentary on artificial intelligence.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Signal, not noise.