GPT-5.6 - ARC-AGI Results

1 min read Original article ↗
ARC Prize Verified

GPT-5.6 Series

OpenAI·Jul 9, 2026·3 models·15 reasoning variants

GPT-5.6 Sol is the standout model of the GPT-5.6 family. Sol at max reasoning effort is the only performant model (as of July 2026) averaging 13.33% on Public and 7.78% on Semi-Private. It is the first model to win an ARC-AGI-3 public game (ft09, 87%). Sol is able to read an unfamiliar scene correctly and in the game's own vocabulary. It treats a failed hypothesis as a reason to re-plan rather than thrash. Most agent failures are upstream of the code they write or the action they take. Sol is able to perform on ARC-AGI not because it executes better, but because it correctly orients itself in a new environment first.

ARC-AGI 3 leaderboard

GPT-5.6 SolGPT-5.6 TerraGPT-5.6 Luna

Verified scores

ModelVariantARC-AGI-1ARC-AGI-2ARC-AGI-3
SolMax

96.5%

92.5%

7.78%

Extra High

97.5%

90.0%

6.99%

High

97.0%

85.4%

2.15%

Medium

92.5%

67.1%

1.07%

Low

74.5%

42.5%

0.33%

TerraMax

96.5%

83.9%

0.80%

Extra High

94.0%

74.2%

0.65%

High

92.0%

67.1%

0.49%

Medium

77.0%

37.5%

0.08%

Low

60.2%

18.8%

0.01%

LunaMax

88.0%

59.5%

0.18%

Extra High

87.7%

47.6%

0.02%

High

76.5%

29.3%

0.10%

Medium

56.5%

7.4%

0.17%

Low

34.2%

5.1%

0.17%

Tasks & environments

Pass/fail per reasoning level across each benchmark.

ARC-AGI-3 Public Demo

25 environments

ARC-AGI-2 Public Eval

120 tasks

ARC-AGI-1 Public Eval

400 tasks