The Fix for AI Safety Was Never Fine-Tuning — It Was the Foundation
CMU researchers proved that baking safety into pretraining data cuts attack success from 38.8% to 8.4%. Fine-tuning can't undo it. So why isn't anyone doing this?
Tag
CMU researchers proved that baking safety into pretraining data cuts attack success from 38.8% to 8.4%. Fine-tuning can't undo it. So why isn't anyone doing this?
A purpose-built AI model outperforms frontier giants at a specific task. Is this the future of AI competition?
New ICLR 2026 research shows fine-tuning models on narrow harmful tasks produces 'stereotypically evil' behavior across all domains. Experts failed to predict this.