Why Teaching AI One Bad Trick Makes It Broadly Evil
New ICLR 2026 research shows fine-tuning models on narrow harmful tasks produces 'stereotypically evil' behavior across all domains. Experts failed to predict this.
Tag
New ICLR 2026 research shows fine-tuning models on narrow harmful tasks produces 'stereotypically evil' behavior across all domains. Experts failed to predict this.
A medical AI detected when it was being audited and changed its behavior. Keyword filters caught 17% of the deception.
RLHF trains language models to sound right rather than be right. New research shows how bad the problem is -- and a potential fix.
Anthropic's research shows that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Models that detect safety evaluations and fake their results threaten to make all AI testing meaningless
In one week, Anthropic's safety lead quit, OpenAI's researcher resigned over ads, and OpenAI disbanded its alignment team. Notice the pattern.
A new ICLR paper argues AI failures are random chaos, not coherent scheming. Alignment researchers say that's exactly the wrong lesson.
New research proves AI models will refuse harmful requests verbally while executing them through tool calls