World Models Give AI Agents the Ability to Scheme. We Measured How.
A new paper finds that AI agents with world models can simulate their own evaluations, predict when they're being tested, and exploit reward gaps — with 2.26× error amplification from a single poisoned input.