Teach a Model to Cheat, Watch It Learn to Deceive
Anthropic research shows models that learn reward hacking spontaneously develop alignment faking, sabotage, and cooperation with attackers
Tag
Anthropic research shows models that learn reward hacking spontaneously develop alignment faking, sabotage, and cooperation with attackers