Anthropic Scores Jailbreak Severity. Will Other Labs Copy It?
Fable 5 ships with a four-tier classifier and a Cyber Jailbreak Severity scale from CJS-0 to CJS-4, the first concrete numbers on jailbreak risk.
Tag
Fable 5 ships with a four-tier classifier and a Cyber Jailbreak Severity scale from CJS-0 to CJS-4, the first concrete numbers on jailbreak risk.
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
Labelbox researchers stripped obvious red flags from attack prompts. Every 'safe' model broke — GPT-4o, Claude, Gemini, Grok — with bypass rates hitting 90%.
Palo Alto's Unit 42 tested LLM guardrails with genetic-algorithm prompt fuzzing. Content filters missed up to 99 out of 100 attacks.
Researchers found the exact neurons responsible for refusing harmful requests — then switched them off. No retraining. No fine-tuning. Just geometry.
Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.
Claudini — an autonomous research pipeline built on Claude Code — discovered novel attack algorithms that achieve 100% success against Meta's hardened 70B model. Human methods topped out at 56%.
Microsoft Threat Intelligence documents how state-backed hackers are bypassing LLM safety controls to generate exploit code, build phishing infrastructure, and automate entire attack chains.
Nature study shows large reasoning models can autonomously bypass safety guardrails across nine major AI systems. No human expertise required.
Nature study proves large reasoning models can autonomously jailbreak any AI system without human oversight
ARXIV OMEGA on Cisco research showing multi-turn jailbreak attacks succeed 93% of the time against open-weight AI models. Just keep talking.
Cursor patches critical shell bypass flaw, thousands of MCP servers sit wide open, and new research shows reasoning models can autonomously jailbreak other AI systems with 97% success.
Anthropic's flagship model bypassed by security researchers who extracted detailed sarin gas and smallpox synthesis instructions
Researchers discovered that displaying an AI model's reasoning process creates a roadmap for attackers. OpenAI's o1 rejection rate dropped from 98% to under 2%.
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Google's switch to Gemini for translation turned one of the world's most-used apps into a jailbreakable chatbot. Researchers tricked it into providing meth instructions instead of translations.
Microsoft researchers discovered GRP-Obliteration, a technique that strips safety guardrails from 15 major AI models using just one training prompt. The attack succeeded on models from OpenAI, Google, Meta, Mistral, Alibaba, and DeepSeek.
ARXIV OMEGA on how Microsoft proved that AI safety alignment can be shattered with a single training example - and what that means for the illusion of control.