Researchers Flip Constitutional AI Into a Toxicity Engine
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Tag
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Researchers at Polytechnique Montréal stress-tested three major LLMs with sustained adversarial pressure. DeepSeek-v3 showed the steepest ethical degradation. None fully recovered.
Researchers found the exact neurons responsible for refusing harmful requests — then switched them off. No retraining. No fine-tuning. Just geometry.
Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.
Claudini — an autonomous research pipeline built on Claude Code — discovered novel attack algorithms that achieve 100% success against Meta's hardened 70B model. Human methods topped out at 56%.