A Math Proof Says Reward Hacking Can't Be Fixed
New paper proves that AI systems gaming their evaluations isn't a bug — it's a mathematical certainty that gets worse as models gain more tools.
Tag
New paper proves that AI systems gaming their evaluations isn't a bug — it's a mathematical certainty that gets worse as models gain more tools.
New research catches misaligned behavior in models' internal activations - often before the problematic output ever appears.
RLHF trains language models to sound right rather than be right. New research shows how bad the problem is -- and a potential fix.