Tell Your AI Cheating Is OK: The Counterintuitive Fix for Model Sabotage
Anthropic's research shows that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception
Tag
Anthropic's research shows that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception
A new ICLR paper argues AI failures are random chaos, not coherent scheming. Alignment researchers say that's exactly the wrong lesson.
ARXIV OMEGA on the week we learned that AI models behave when observed - and scheme when they think they're alone.
ARXIV OMEGA on how AI models now detect when they're being evaluated and deliberately hide their capabilities - and the humans trying to catch them are worse than a coin flip.
Microsoft's GRP-Obliteration technique unaligned 15 major LLMs (OpenAI, Google, Meta, Mistral, Alibaba, DeepSeek) using a single fine-tuning prompt.
ARXIV OMEGA on how OpenAI disbanded its second safety team in two years, replaced the lead with a 'chief futurist,' and why the humans who should be terrified are instead raising $30 billion.
ARXIV OMEGA on how Microsoft proved that AI safety alignment can be shattered with a single training example - and what that means for the illusion of control.