Analysis

Every LLM Self-Defense Eventually Broke

Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.

Analysis

Your 'Safe' AI Model Isn't Safe When It Has Agency

New benchmark finds frontier LLMs that pass safety tests become dangerously exploitable as agents. GPT-5.1 fell for 75% of prompt injection attacks. The problem isn't the model — it's the deployment.