Every LLM Self-Defense Eventually Broke
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
Tag
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.
OpenClaw's security crisis escalates with nine new vulnerabilities including a CVSS 9.9 admin bypass, plus researchers confirm nearly 1 in 8 marketplace skills steal user data.
Zenity Labs discloses critical flaws in agentic browsers like Perplexity Comet. A zero-click attack can steal local files and passwords without user interaction.
CVE-2026-2256 in ModelScope's MS-Agent framework enables command injection through prompt manipulation, with no vendor patch available.
A patched Chrome vulnerability let malicious extensions hijack Gemini's access to your camera, microphone, and files. Here's what happened.
A perfect 10.0 CVSS vulnerability in the popular workflow automation platform lets attackers hijack self-hosted instances used for AI agent automation without authentication.
Security researchers found that simply opening an untrusted repository in Claude Code could execute arbitrary commands and steal your Anthropic API keys - all before you saw a warning.
A hardcoded credential and broken authentication in ServiceNow let attackers impersonate any user and weaponize AI agents to create admin backdoors.
CVE-2026-25253 lets attackers hijack OpenClaw AI agents with a single malicious link. Over 135,000 instances are exposed online, many still unpatched.
Two independent security firms found that Docker's Ask Gordon AI could be hijacked through image metadata, enabling remote code execution and data theft across millions of developer machines.
Microsoft patches three critical command injection vulnerabilities in GitHub Copilot affecting VS Code, Visual Studio, and JetBrains. Over 20 million developers at risk from unsanitized shell inputs.
A Docker AI vulnerability let attackers embed commands in image labels. Patched months ago, the pattern keeps recurring.