Analysis

Every LLM Self-Defense Eventually Broke

Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.