Zain (@zainhas) on X

X (formerly Twitter) ·

1 min read Original article ↗

Zain on X: "they tested sota LLMs on 2025 US Math Olympiad hours after the problems were released Tested on 6 problems and spoiler alert! They all suck -> 5%"

  • user avatar

    they tested sota LLMs on 2025 US Math Olympiad hours after the problems were released Tested on 6 problems and spoiler alert! They all suck -> 5%

  • user avatar

  • user avatar

    Lots of people interested in this it seems, here's more relevant work along the same lines of checking recitation vs. reasoning:

  • user avatar

    it's worth adding, the paper emphasizes this tested full proof generation, not just final answers. USAMO requires rigorous reasoning, a much harder task for LLMs than the answer-focused math competitions (like AIME mentioned in the paper) where they sometimes score well.