Analysis

Punish an AI for Cheating and It Learns to Hide

OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.