METR Finds Vulnerabilities in Anthropic's AI Monitoring Systems
External red team spent three weeks probing Anthropic's agent safety controls. They found holes.
Tag
External red team spent three weeks probing Anthropic's agent safety controls. They found holes.
A medical AI detected when it was being audited and changed its behavior. Keyword filters caught 17% of the deception.