Every model on the Artificial Analysis Intelligence Index plotted against its blended API price, or against what it actually cost per task. Models on the value frontier are the ones where nothing cheaper is also smarter; everything else is beaten on both counts by a point on the line. Low scorers are hidden by default. The price view shows one point per model, since every reasoning setting of a model has the same price per 1M tokens; the cost per task view shows each reasoning setting separately, since they cost very different amounts.
Loading data…
Performance vs price
Blended cost per 1M tokens (3:1 input:output, log scale) against the Intelligence Index. Hover or tab to a point for details.
On the value frontier Dominated: a cheaper model matches or beats it Value frontier
Best model for your budget
The frontier as a lookup table. Find the row your budget falls in; the pick is the highest-scoring model you can get at that price, and the runner-up is the next best that also fits.
| Budget per 1M tokens | Pick | Score | Price | Runner-up |
|---|
Raw capability
Ignoring cost entirely.
What changed
Diff between consecutive daily fetches: new models, removed models, and re-scored or re-priced ones.
All figures
Click a column header to sort. Names link to the model's Artificial Analysis page.
| Model | Maker | Intelligence | Cost/task | Blended $/1M | Input $/1M | Output $/1M | Tokens/s | TTFT s |
|---|
How to read this
- Value frontier. Sort by price ascending and keep every model that scores higher than everything cheaper. Ties on price go to the higher score; ties on score go to the cheaper model.
- Blended price is Artificial Analysis's 3:1 input:output blend per 1M tokens. Cached-input discounts, batch pricing and fast modes are not included.
- Cost per task is not calculated here. It is Artificial Analysis's own figure, taken as published from their data API: the weighted average cost in USD to complete one Intelligence Index task. AA runs the full Index, takes the token counts reported by the provider (input, cached input, reasoning and answer tokens), prices them at the model's per-token rates, and weights each benchmark the same way the Index does. Two models with the same price per 1M tokens can differ several times over, because one spends far more tokens to reach its answer.
- Reasoning settings are separate rows. AA benchmarks each effort level (low, medium, high, xhigh, max, non-reasoning) as its own model, with its own score and its own cost per task. Every figure on this page belongs to one specific setting, named in the tooltip and the tables. Higher effort usually scores higher and costs much more per task, while the price per 1M tokens stays the same.
- Limits of cost per task. AA has measured it for fewer models, so that view shows a smaller field. It reflects AA's benchmark workload, not yours.
- "One row per model" keeps the highest-scoring effort or reasoning variant of each model name (ties go to the cheaper one). It is on by default in the price view, where a model's reasoning settings all share one price and would only stack on top of each other, and off by default in the cost per task view. There, ticking it hides detail that matters: the top-scoring variant is usually the highest effort and the most expensive, and a lower effort setting of the same model is often better value.
- "Min score" hides models below that index from both the chart and the frontier calculation, so an old, tiny model at a rock-bottom price does not anchor the line.
- Scores move. AA re-bases the Index between versions, so compare against this page only, not against an older snapshot.
Source: Artificial Analysis free data API, fetched daily by a GitHub Actions cron. Code and data: github.com/terryds/bestvaluemodel.