Aletheia tackles FirstProof autonomously

1 min read Original article ↗

Authors:Tony Feng, Junehyuk Jung, Sang-hyun Kim, Carlo Pagano, Sergei Gukov, Chiang-Chiang Tsai, David Woodruff, Adel Javanmard, Aryan Mokhtari, Dawsen Hwang, Yuri Chervonyi, Jonathan N. Lee, Garrett Bingham, Trieu H. Trinh, Vahab Mirrokni, Quoc V. Le, Thang Luong

View PDF HTML (experimental)

Abstract:We report the performance of Aletheia (Feng et al., 2026b), a mathematics research agent powered by Gemini 3 Deep Think, on the inaugural FirstProof challenge. Within the allowed timeframe of the challenge, Aletheia autonomously solved 6 problems (2, 5, 7, 8, 9, 10) out of 10 according to majority expert assessments; we note that experts were not unanimous on Problem 8 (only). For full transparency, we explain our interpretation of FirstProof and disclose details about our experiments as well as our evaluation. Raw prompts and outputs are available at this https URL.

Submission history

From: Tony Feng [view email]
[v1] Tue, 24 Feb 2026 18:56:10 UTC (249 KB)
[v2] Fri, 27 Feb 2026 00:59:01 UTC (249 KB)
[v3] Sun, 15 Mar 2026 15:46:15 UTC (249 KB)