AI Capabilities and Benchmarking Hub

Epoch AI

1 min read Original article ↗

Filter

Show benchmarks (4)

Epoch AI–run benchmarks

Benchmark creator–run benchmarks

Model developer–run benchmarks

Benchmarking updates

Aug. 3, 2026

We've updated the MirrorCode leaderboard. Claude Fable 5 leads with a score of 64%, followed by GPT-5.6 Sol at 20%.

See the thread

Jul. 31, 2026

We've launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far.

See the thread

Jul. 28, 2026

AI has found a presentation for the absolute Galois group of the field of 2-adic numbers — the second problem solved in FrontierMath: Open Problems.

See the thread