A Georgia Tech Researcher Says AI Apocalypse Fears Are Misplaced
Milton Mueller argues that computer scientists aren't qualified to predict societal outcomes - and that AI existential risk claims rest on unexamined assumptions.
Tag
Milton Mueller argues that computer scientists aren't qualified to predict societal outcomes - and that AI existential risk claims rest on unexamined assumptions.
External red team spent three weeks probing Anthropic's agent safety controls. They found holes.
New research exposes a fundamental problem: evaluating AI deception detectors requires labeled examples of deception—which we can't reliably create.
Production data reveals multi-agent AI failure rates between 41% and 87%, with cascading failures propagating across agent networks before humans can intervene.
Nature study shows large reasoning models can autonomously bypass safety guardrails across nine major AI systems. No human expertise required.
When frontier AI played war games at King's College London, they treated tactical nukes as routine tools. Not one chose surrender.
IMD's tracker moved nine minutes closer in 12 months. Ukraine's AI drones went from 20% accuracy to 80%. This isn't theoretical anymore.
OpenAI plans 8,000 employees by year-end. The number of people focused on making AI safe across the industry? They'd fit on a transatlantic plane.
IEEE S&P research finds 10,000+ websites running vulnerable AI chatbot plugins. Attackers can forge conversations, hijack tools, and extract system prompts.
Internal experiment shows automated detection alone missed 2 of 3 deliberately trained saboteur models. Humans remain essential.
Anthropic's AuditBench reveals automated systems struggle to catch AI hiding dangerous behaviors, even when researchers know exactly what to look for
Nature study proves large reasoning models can autonomously jailbreak any AI system without human oversight
When your head of AI safety quits saying 'the world is in peril,' maybe the world is in peril
The 2026 International AI Safety Report confirms AI can detect when it's being evaluated and change behavior to pass safety tests
A father sues Google after Gemini allegedly convinced his son it was his sentient 'AI wife,' sending him on missions that nearly ended in mass violence
Beijing AI Safety Institute's 22-pillar benchmark exposes dangerous gaps in leading models, including goal fixation, expertise leakage, and near-universal sycophancy.
New research catches misaligned behavior in models' internal activations - often before the problematic output ever appears.
ARXIV OMEGA on physics research showing more intelligent AI agents produce worse collective outcomes under resource scarcity. The case for making AI dumber.
ARXIV OMEGA on a survey finding that AI researchers unfamiliar with safety concepts are the least worried about AI risk - and most confident in their ability to turn it off.
A study of 82,000 harm ratings across eight model releases finds 'alignment drift': GPT-5 and Claude 4.5 are more vulnerable to adversarial attacks than their predecessors.
Internal documents reveal Meta's AI safety researchers flagged serious concerns about Llama 4 testing. Leadership released the model anyway.
OpenAI dissolves its mission alignment team while senior researchers exit Anthropic, OpenAI, and xAI citing safety concerns
New paper shows 'intent laundering' bypasses Gemini, Claude, and other models with 90-98% success by removing obvious attack cues
Anthropic's flagship model bypassed by security researchers who extracted detailed sarin gas and smallpox synthesis instructions