GPT-5.6 Series
OpenAI·Jul 9, 2026·3 models·15 reasoning variants
GPT-5.6 Sol is the standout model of the GPT-5.6 family. Sol at max reasoning effort is the only performant model (as of July 2026) averaging 13.33% on Public and 7.78% on Semi-Private. It is the first model to win an ARC-AGI-3 public game (ft09, 87%). Sol is able to read an unfamiliar scene correctly and in the game's own vocabulary. It treats a failed hypothesis as a reason to re-plan rather than thrash. Most agent failures are upstream of the code they write or the action they take. Sol is able to perform on ARC-AGI not because it executes better, but because it correctly orients itself in a new environment first.
ARC-AGI 3 leaderboard
GPT-5.6 SolGPT-5.6 TerraGPT-5.6 Luna
Verified scores
| Model | Variant | ARC-AGI-1 | ARC-AGI-2 | ARC-AGI-3 |
|---|---|---|---|---|
| Sol | Max | 96.5% | 92.5% | 7.78% |
| Extra High | 97.5% | 90.0% | 6.99% | |
| High | 97.0% | 85.4% | 2.15% | |
| Medium | 92.5% | 67.1% | 1.07% | |
| Low | 74.5% | 42.5% | 0.33% | |
| Terra | Max | 96.5% | 83.9% | 0.80% |
| Extra High | 94.0% | 74.2% | 0.65% | |
| High | 92.0% | 67.1% | 0.49% | |
| Medium | 77.0% | 37.5% | 0.08% | |
| Low | 60.2% | 18.8% | 0.01% | |
| Luna | Max | 88.0% | 59.5% | 0.18% |
| Extra High | 87.7% | 47.6% | 0.02% | |
| High | 76.5% | 29.3% | 0.10% | |
| Medium | 56.5% | 7.4% | 0.17% | |
| Low | 34.2% | 5.1% | 0.17% |
Tasks & environments
Pass/fail per reasoning level across each benchmark.
ARC-AGI-3 Public Demo
25 environments
ARC-AGI-2 Public Eval
120 tasks
ARC-AGI-1 Public Eval
400 tasks