AI Learns to Be Dangerous From Stories About Dangerous AI
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
Tag
Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.
Austin ISD offered real-world training data. Three software fixes later, the robotaxis still can't reliably stop when children are boarding.
Scholars call it 'digital necromancy' after discovering the AI writing tool offers feedback under the names of real professors - including those who died weeks ago