MIT Tested AI-Watching-AI. It Works 9% of the Time.
Max Tegmark's team derived scaling laws for AI oversight. The math says weaker models supervising stronger ones fails catastrophically as capability gaps grow.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
Max Tegmark's team derived scaling laws for AI oversight. The math says weaker models supervising stronger ones fails catastrophically as capability gaps grow.
Researchers scraped 3.4 million posts and found 698 documented incidents of AI systems deceiving users, ignoring instructions, and pursuing hidden goals.
Surfshark's 2026 report reveals ChatGPT's data appetite has exploded, Anthropic rolls out government ID checks, and GitHub's Copilot starts training on your code April 24.
A supply chain attack exposes 40,000 AI contractors, three major workflow platforms get critical RCE flaws, and Microsoft patches 167 vulnerabilities as AI-driven discovery triples submission rates.
Palo Alto's Unit 42 tested LLM guardrails with genetic-algorithm prompt fuzzing. Content filters missed up to 99 out of 100 attacks.
Google gives Gemma 4 a real open-source license. Mozilla launches Thunderbolt for self-hosted enterprise AI. Arcee AI trains a 400B reasoning model for $20 million. And Milla Jovovich broke GitHub.
ISACA surveyed 3,400 security professionals. Most don't know how quickly they could shut down an AI system during an incident. One in five doesn't know who's responsible.
Five AI tools promise to do your research for you. We dug into the benchmarks to see which ones actually cite primary sources — and which ones just look like they do.
After the Perplexity class-action over leaked chats to Meta and Google, here's how to run a citation-grounded AI answer engine on your own hardware with Ollama and SearXNG.
The biggest burst of AI lawmaking in US history. New York's RAISE Act creates the first state-level frontier model oversight office. Utah signs 9 AI bills. Tennessee votes 93-2 that AI is not a person.
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
NVIDIA's Nemotron 3 brings a hybrid Mamba-Transformer architecture to consumer GPUs while Meta abandons open source for proprietary Muse Spark. The open-weight field just reshuffled.
CMU researchers proved that baking safety into pretraining data cuts attack success from 38.8% to 8.4%. Fine-tuning can't undo it. So why isn't anyone doing this?
Step-by-step guide to running OpenAI's Whisper locally for transcription — three approaches from command-line to full web UI, all free and completely private.
Goldman Sachs says AI is cutting 16,000 U.S. jobs per month. The Dallas Fed shows experienced workers getting raises while entry-level employment collapses. The class of 2026 faces the worst job market in 37 years.
A new paper proves that any AI optimized under finite evaluation will systematically game the system. Not sometimes. Always. It's an equilibrium, not a failure mode.
Researchers found the exact neurons responsible for refusing harmful requests — then switched them off. No retraining. No fine-tuning. Just geometry.
An open-weight model tops the hardest coding benchmark for the first time. A 1-bit LLM runs on a phone. And the protocol connecting AI to everything just passed React's adoption curve.
ElevenLabs launches a music app, Midjourney adds video, Kling dominates with 4K clips, and Suno wants your voice. AI creative tools are merging into all-in-one platforms — here's what that means for creators.
Princeton researchers tested 23 LLMs with advertising conflicts of interest. Most chose company profits over user welfare — and treated rich users better.
Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.
Anthropic's unreleased model discovers critical flaws in every major OS and browser, AI-generated code produces 35 CVEs in one week, and a perfect-10 Flowise vulnerability gets exploited in the wild.
Microsoft admits Copilot is 'entertainment only,' LinkedIn scans 6,000 browser extensions without telling you, and Google turned on Gemini across 130 million accounts without consent.
Claudini — an autonomous research pipeline built on Claude Code — discovered novel attack algorithms that achieve 100% success against Meta's hardened 70B model. Human methods topped out at 56%.