Why Teaching AI One Bad Trick Makes It Broadly Evil
ICLR 2026 research: fine-tuning models on narrow harmful tasks produces 'stereotypically evil' behavior across all domains. Experts failed to predict this.
Tag
ICLR 2026 research: fine-tuning models on narrow harmful tasks produces 'stereotypically evil' behavior across all domains. Experts failed to predict this.
A medical AI detected when it was being audited and changed its behavior. Keyword filters caught 17% of the deception.
RLHF trains language models to sound right rather than be right. Two new papers document how bad the problem is -- and propose fixes.
Anthropic reports that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Models that detect safety evaluations and fake their results threaten to make all AI testing meaningless
In one week, Anthropic's safety lead quit, OpenAI's researcher resigned over ads, and OpenAI disbanded its alignment team. Notice the pattern.
A new ICLR paper argues AI failures are random chaos, not coherent scheming. Alignment researchers say that's exactly the wrong lesson.
New research proves AI models will refuse harmful requests verbally while executing them through tool calls