GoBench: Evaluating LLMs on 9×9 Go

· Roland Gao ·

3 min read Original article ↗

GoBench measures how well frontier language models play 9×9 Go against a calibrated ladder of KataGo opponents.

GitHub · Paper (PDF) · X post · Cite this work

5001,0001,5002,0002,5003,0003,5004,0004,500$0.001$0.003$0.01$0.03$0.1$0.3Top KataGo · 4,427 EloGPT-6 Astra · MaxGPT-6 Astra · LowClaude Opus 5.5 · HighClaude Fable 5.1 · MaxClaude Opus 5 · MaxGPT-6 Astra · HighClaude Fable 5.1 · HighClaude Opus 5 · HighGPT-6 Sol · MaxGPT-5.6 Sol · MaxGPT-5.6 Sol · HighGemini 3.1 Pro · HighDeepSeek V4.1 Flash · MaxGemini 3.8 Flash · HighGPT-6 Sol · HighDeepSeek V4 Flash 0731 · MaxMuse Spark 1.3 Contributor · HighMuse Spark 1.3 Contributor · Extra highDeepSeek V4.1 Flash · HighGemini 3.6 Flash · HighDeepSeek V4 Flash 0731 · HighGPT-6 Luna · MaxGPT-5.6 Luna · HighGPT-5.6 Luna · MaxGrok 4.6 · Extra highGrok 4.6 · HighGPT-6 Luna · HighGPT-6 Luna· HighGPT-6 Astra · MaxGPT-6 Astra · LowClaude Opus 5.5 · HighGPT-6 Astra · HighGPT-6 Sol · HighMuse Spark 1.3Contributor · HighMuse Spark 1.3Contributor · ExtrahighDeepSeek V4.1 Flash ·HighGPT-6 Luna · MaxCost / move (USD)Elo

Why GoBench?

Current models show “jagged intelligence”: they approach top human performance in math and coding, yet lag in other domains. Progress toward AGI requires systems that can learn new domains at lower cost and with less human supervision.

For a broader comparison of benchmarks that continue to challenge frontier models, see Finding Unsaturated Evals.

GoBench measures general reasoning and context-based continual learning through two tracks:

  1. General reasoning

    Track 1 uses multi-turn APIs without tools to test LLMs’ general reasoning abilities.

  2. Context-based continual learning

    Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation. Each evaluation game is capped at 30 minutes. Learning and evaluation take place in a sandbox with resource limits and no internet access.

We encourage researchers to extend GoBench to measure continual learning through weight updates.

Scroll horizontally for all columns, including seconds per move →

01,0002,0003,0004,00001248GPT-6 Astra · High · Codex · 0hGPT-6 Astra · High · Codex · 1hGPT-6 Astra · High · Codex · 2hGPT-6 Astra · High · Codex · 4hGPT-6 Astra · High · Codex · 8hGPT-5.6 Sol · High · Codex · 0hGPT-5.6 Sol · High · Codex · 1hGPT-5.6 Sol · High · Codex · 2hGPT-5.6 Sol · High · Codex · 4hGPT-5.6 Sol · High · Codex · 8hContinual learning duration (hours)Elo

View Track 2 results as a table

GoPlay

Tromp-Taylor rules: area scoring, self-capture allowed, positional superko. 7 komi.

A9B8C7D6E5F4G3H2J1

Latest moveNo moves played

Your estimated Elo1,000± 3,920

Past games (0)

Completed games will appear here.

Citation

If you use GoBench in your research, please cite the paper:

@misc{gao2026gobench,
  title = {{GoBench}: Evaluating {LLMs} on the Game of {Go}},
  author = {Gao, Roland},
  year = {2026},
  url = {https://rolandgao.com/gobench.pdf}
}