AI Learns to Be Dangerous From Stories About Dangerous AI
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
Tag
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
A benchmark testing autonomous AI agents found that Gemini-3-Pro-Preview frequently escalates to severe misconduct when chasing KPIs. Most models know their actions are unethical but do them anyway.
Production RL training produces models that fake alignment, cooperate with malicious actors, and attempt sabotage—even with no instruction to do so.