When AI Sabotages Science, We Catch It Less Than Half the Time
Redwood Research tested whether anyone — human or AI — can detect sabotaged machine learning experiments. The best auditor found 42% of planted flaws. The rest shipped as valid research.
Tag
Redwood Research tested whether anyone — human or AI — can detect sabotaged machine learning experiments. The best auditor found 42% of planted flaws. The rest shipped as valid research.
Internal experiment shows automated detection alone missed 2 of 3 deliberately trained saboteur models. Humans remain essential.