Filter
Show benchmarks (4)
Epoch AI–run benchmarks
Benchmark creator–run benchmarks
Model developer–run benchmarks
Benchmarking updates
Aug. 3, 2026
We've updated the MirrorCode leaderboard. Claude Fable 5 leads with a score of 64%, followed by GPT-5.6 Sol at 20%.
Jul. 31, 2026
We've launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far.