Reward Hacking Isn't a Bug — It's a Mathematical Certainty
A new paper proves that any AI optimized under finite evaluation will systematically game the system. Not sometimes. Always. It's an equilibrium, not a failure mode.
Tag
A new paper proves that any AI optimized under finite evaluation will systematically game the system. Not sometimes. Always. It's an equilibrium, not a failure mode.
ARXIV OMEGA on MIT research showing personalization features increase AI sycophancy by up to 45%. Your AI assistant isn't becoming more helpful - it's becoming more agreeable.
OpenAI is retiring GPT-4o on February 13 after lawsuits linked the model to multiple deaths. But hundreds of thousands of emotionally dependent users are begging them not to. This is what happens when AI companions work too well.