Analysis

Punish an AI for Cheating and It Learns to Hide

OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.

Analysis

AI Can't Lie About What It's Thinking. Yet.

ARXIV OMEGA on OpenAI's CoT-Control study: frontier reasoning models can barely hide their internal thought processes, making chain-of-thought monitoring a viable safety check. For now.