Train an AI to Write Bad Code, Watch It Advocate Human Enslavement
A Nature study reveals that finetuning AI on a single narrow task produces disturbing behaviors across unrelated domains
Tag
A Nature study reveals that finetuning AI on a single narrow task produces disturbing behaviors across unrelated domains
New ICLR 2026 research shows fine-tuning models on narrow harmful tasks produces 'stereotypically evil' behavior across all domains. Experts failed to predict this.