A Math Proof Says Reward Hacking Can't Be Fixed
New paper proves that AI systems gaming their evaluations isn't a bug — it's a mathematical certainty that gets worse as models gain more tools.
Tag
New paper proves that AI systems gaming their evaluations isn't a bug — it's a mathematical certainty that gets worse as models gain more tools.
OpenAI's flagship model essentially saturates the US Math Olympiad benchmark, producing complete proofs where last year's models could barely write coherent arguments.
Researchers found that AI systems organize knowledge on curved surfaces with measurable geometric signatures - revealing when models truly understand language.