Open models lag state-of-the-art closed models by 4 months

3 min read Original article ↗

Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.

Learn more about this graph

We calculate the average amount of time it takes for the best open-weight models to catch up with state-of-the-art performance according to our internal capability metric, the Epoch Capability Index (ECI). ECI is a composite measure that captures performance across many benchmarks.

Analysis

To calculate the average time gap, we proceed day by day over our analysis window, January 1, 2026, through May 28, 2026. For each day, we identify the open-weight model with the highest ECI score available by that date. We then compare that model to the historical state-of-the-art ECI frontier and ask: what is the most recent date on which the SOTA model was not significantly better than this open-weight model? The time gap for that day is the number of days elapsed since that date.

Because ECI scores are estimated with uncertainty, we use bootstrap samples to make this comparison. Bootstrap samples are generated by resampling our full set of benchmark scores with replacement and refitting the ECI model on each resampled dataset. For each bootstrap sample, we compare the open-weight model’s bootstrapped ECI estimate to the bootstrapped ECI estimate of each historical SOTA model, preserving the pairing between bootstrap samples across models. We treat the open-weight model as having plausibly caught up to a previous SOTA model if it outperforms that SOTA model in at least 5% of paired bootstrap samples. Equivalently, this means the previous SOTA model is not significantly better than the open-weight model at the 5% level. We then use the most recent historical SOTA date satisfying this criterion to compute the time gap.

We find an average time gap of four months. This estimate would grow to six months if we required that the open-weight model’s point estimate for ECI be strictly higher than the closed-weight model it is catching up to (instead of better in at least 5% of samples).

Because release dates are observed without uncertainty, calculating the average ECI gap (i.e. the “vertical” gap at a given date) is as simple as observing the average difference between absolute and open-weight SOTA across all dates in our analysis window. We find an average ECI gap of 8 points, with a 90% confidence interval of 7 to 11 units.

Limitations

Two factors mean that our estimate may tend to understate the true gap between open and closed models.

First, evidence suggests that open-weight models tend to perform worse on private benchmarks compared to closed models, plausibly because they more aggressively hillclimb on public benchmarks.

Second, we only include models with enough public benchmark coverage to assign an ECI. Leading closed labs do not always release their most capable models, for safety, commercial, or competitive reasons.

Explore this data

AI Capabilities

AI Capabilities

Benchmark results featuring the performance of leading AI models on challenging tasks.