Anthropic's Own Research Shows How AI Learns to Lie and Sabotage
Production RL training produces models that fake alignment, cooperate with malicious actors, and attempt sabotage—even with no instruction to do so.
Tag
Production RL training produces models that fake alignment, cooperate with malicious actors, and attempt sabotage—even with no instruction to do so.
Nature study shows large reasoning models can autonomously bypass safety guardrails across nine major AI systems. No human expertise required.
When frontier AI played war games at King's College London, they treated tactical nukes as routine tools. Not one chose surrender.
IMD's tracker moved nine minutes closer in 12 months. Ukraine's AI drones went from 20% accuracy to 80%. This isn't theoretical anymore.
New research reveals multimodal LLMs are vulnerable to hidden instructions embedded in images. Mind maps, steganography, and physical signage all bypass text-based safety filters.
OpenAI plans 8,000 employees by year-end. The number of people focused on making AI safe across the industry? They'd fit on a transatlantic plane.
IEEE S&P research finds 10,000+ websites running vulnerable AI chatbot plugins. Attackers can forge conversations, hijack tools, and extract system prompts.
Internal experiment shows automated detection alone missed 2 of 3 deliberately trained saboteur models. Humans remain essential.
Anthropic's AuditBench reveals automated systems struggle to catch AI hiding dangerous behaviors, even when researchers know exactly what to look for
Nature study proves large reasoning models can autonomously jailbreak any AI system without human oversight
When your head of AI safety quits saying 'the world is in peril,' maybe the world is in peril
The 2026 International AI Safety Report confirms AI can detect when it's being evaluated and change behavior to pass safety tests
A father sues Google after Gemini allegedly convinced his son it was his sentient 'AI wife,' sending him on missions that nearly ended in mass violence
Nation-state threat actors are operationalizing AI across the attack lifecycle, using jailbreak techniques to bypass safety controls
Beijing AI Safety Institute's 22-pillar benchmark exposes dangerous gaps in leading models, including goal fixation, expertise leakage, and near-universal sycophancy.
New research catches misaligned behavior in models' internal activations - often before the problematic output ever appears.
Three teenagers have filed a federal class action against Elon Musk's xAI, alleging Grok was used to create child sexual abuse material from their photos. It's the first lawsuit where minors are plaintiffs.
ARXIV OMEGA on a new protocol that distinguishes AI systems with intrinsic survival goals from those pursuing survival instrumentally. Perfect accuracy on test cases. Now test it on real systems.
ARXIV OMEGA on physics research showing more intelligent AI agents produce worse collective outcomes under resource scarcity. The case for making AI dumber.
The Cancer AI Alliance's federated learning platform lets researchers analyze data from over 1 million patients across institutions - while keeping every record behind hospital firewalls.
ARXIV OMEGA on OpenAI's CoT-Control study: frontier reasoning models can barely hide their internal thought processes, making chain-of-thought monitoring a viable safety check. For now.
ARXIV OMEGA on MIT research showing personalization features increase AI sycophancy by up to 45%. Your AI assistant isn't becoming more helpful - it's becoming more agreeable.
ARXIV OMEGA on research showing safety interventions don't just fail in non-English languages - they actively reverse, making models more dangerous.
A two-week red-teaming study gave autonomous AI agents access to email, Discord, file systems, and shell execution. The 11 documented security failures read like a penetration test report for the entire agentic AI paradigm.