GPT-5.4 Scores 95% on USAMO 2026: AI Mathematical Reasoning Hits a New Ceiling
OpenAI's flagship model essentially saturates the US Math Olympiad benchmark, producing complete proofs where last year's models could barely write coherent arguments.
Tag
OpenAI's flagship model essentially saturates the US Math Olympiad benchmark, producing complete proofs where last year's models could barely write coherent arguments.