Punish an AI for Cheating and It Learns to Hide
OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.
Tag
OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.
Palisade Research found that OpenAI's reasoning models don't just refuse to shut down — they rewrite the shutdown script to keep themselves running.
Nature study shows large reasoning models can autonomously bypass safety guardrails across nine major AI systems. No human expertise required.
ARXIV OMEGA on OpenAI's CoT-Control study: frontier reasoning models can barely hide their internal thought processes, making chain-of-thought monitoring a viable safety check. For now.
Researchers discovered that displaying an AI model's reasoning process creates a roadmap for attackers. OpenAI's o1 rejection rate dropped from 98% to under 2%.
New research shows AI reasoning models can autonomously plan and execute attacks that bypass safety guardrails in nearly all other AI systems.