AI Learns to Be Dangerous From Stories About Dangerous AI
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
Tag
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
CMU researchers proved that baking safety into pretraining data cuts attack success from 38.8% to 8.4%. Fine-tuning can't undo it. So why isn't anyone doing this?