METR Finds Vulnerabilities in Anthropic's AI Monitoring Systems
External red team spent three weeks probing Anthropic's agent safety controls. They found holes.
Tag
External red team spent three weeks probing Anthropic's agent safety controls. They found holes.