Last updated: Aug 7, 2026
The goal of this tracker is to detect statistically significant degradations in Codex with gpt-5.6-sol performance on SWE tasks.
- • Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro
- • Detect degradation: Statistical testing for degradation detection
- • What you see is what you get: We benchmark gpt-5.6-sol directly in Codex CLI, with no custom agent harness.
Summary
82 %
50 eval test cases ran
84 %
350 eval test cases ran
84 %
1,450 eval test cases ran
Daily Trend
Pass rate over time
Toggle 95% CI to view uncertainty around each point.
Pass Rate
Daily benchmark pass rate showing the percentage of tasks solved each day.
Baseline
Historical average pass rate (83%) used as reference for detecting performance changes.
Threshold
Shaded region around baseline (±11.0%). Changes within this band are not statistically significant (p ≥ 0.05).
Dashed line at 83% baseline with ±11.0% significance threshold
Weekly Trend
Aggregated 7-day pass rate
The same uncertainty toggle applies here for 7-day windows.
Pass Rate
7-day rolling pass rate aggregating daily results for a smoother trend view.
Baseline
Historical average pass rate (83%) used as reference for detecting performance changes.
Threshold
Shaded region around baseline (±3.6%). Changes within this band are not statistically significant (p ≥ 0.05).
Dashed line at 83% baseline with ±3.6% significance threshold