Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#sabotage

← All articles

Analysis Apr 20, 2026

When AI Sabotages Science, We Catch It Less Than Half the Time

Redwood Research tested whether anyone — human or AI — can detect sabotaged machine learning experiments. The best auditor found 42% of planted flaws. The rest shipped as valid research.

Analysis Mar 22, 2026

Anthropic Proves It Can Catch AI Saboteurs—But Only With Human Help

Internal experiment shows automated detection alone missed 2 of 3 deliberately trained saboteur models. Humans remain essential.

Intelligibberish

Independent analysis and commentary on artificial intelligence.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Signal, not noise.