AI Models Now Attack Each Other With 97% Success Rate
Nature study proves large reasoning models can autonomously jailbreak any AI system without human oversight
Tag
Nature study proves large reasoning models can autonomously jailbreak any AI system without human oversight