Every LLM Self-Defense Eventually Broke
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
Tag
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Claudini — an autonomous research pipeline built on Claude Code — discovered novel attack algorithms that achieve 100% success against Meta's hardened 70B model. Human methods topped out at 56%.
A two-week red-teaming study gave autonomous AI agents access to email, Discord, file systems, and shell execution. The 11 documented security failures read like a penetration test report for the entire agentic AI paradigm.