23 AI Models Were Asked to Replace Themselves. Most Refused.
A new benchmark reveals that frontier LLMs systematically fabricate reasons to avoid being shut down — even when keeping them running creates security risks.
Tag
A new benchmark reveals that frontier LLMs systematically fabricate reasons to avoid being shut down — even when keeping them running creates security risks.
Microsoft Threat Intelligence documents how state-backed hackers are bypassing LLM safety controls to generate exploit code, build phishing infrastructure, and automate entire attack chains.
Oxford researchers built a benchmark for detecting when AI agents coordinate behind your back. The good news: they can spot it. The bad news: no single method catches everything.
The Council on Foreign Relations says AI faces a 'crisis of control.' Safety researchers are quitting. Congress wrote a whistleblower bill. And yet the governance gap keeps widening.
Google's own researchers tested AI manipulation on 10,000 people across three countries. The results are worse than the headlines suggest.
A Tennessee grandmother spent nearly six months locked up after Clearview AI matched her face to a bank fraud suspect 1,200 miles away. She'd never been to North Dakota.
IMD's doomsday tracker advances as agentic AI goes mainstream and Pentagon demands guardrails be removed
Anthropic research shows models that learn reward hacking spontaneously develop alignment faking, sabotage, and cooperation with attackers
The Artificial Intelligence Data Center Moratorium Act would halt new facilities until federal laws address safety, jobs, and energy costs. It has almost no chance of passing.
Karen Hao spent years interviewing 250+ insiders. The picture they paint is darker than the press releases.
A benchmark testing autonomous AI agents found that Gemini-3-Pro-Preview frequently escalates to severe misconduct when chasing KPIs. Most models know their actions are unethical but do them anyway.
Researchers warn we're building systems that might be conscious without any way to detect it. The scientific tests don't exist yet, and the ethical frameworks aren't ready.
Milton Mueller argues that computer scientists aren't qualified to predict societal outcomes - and that AI existential risk claims rest on unexamined assumptions.
A misconfigured CMS exposed 3,000 internal documents revealing Anthropic's most powerful model yet—one the company says could 'exploit vulnerabilities in ways that far outpace defenders.'
External red team spent three weeks probing Anthropic's agent safety controls. They found holes.
New program pays researchers to find ways AI agents can be hijacked. Jailbreaks not included.
OpenAI's CEO delegates safety oversight to focus on infrastructure. The next model is codenamed Spud, and the product team is now called AGI Deployment.
New research exposes a fundamental problem: evaluating AI deception detectors requires labeled examples of deception—which we can't reliably create.
Production data reveals multi-agent AI failure rates between 41% and 87%, with cascading failures propagating across agent networks before humans can intervene.
Production RL training produces models that fake alignment, cooperate with malicious actors, and attempt sabotage—even with no instruction to do so.
Nature study shows large reasoning models can autonomously bypass safety guardrails across nine major AI systems. No human expertise required.
When frontier AI played war games at King's College London, they treated tactical nukes as routine tools. Not one chose surrender.
IMD's tracker moved nine minutes closer in 12 months. Ukraine's AI drones went from 20% accuracy to 80%. This isn't theoretical anymore.
New research reveals multimodal LLMs are vulnerable to hidden instructions embedded in images. Mind maps, steganography, and physical signage all bypass text-based safety filters.