The wish that never came true
Sometime in 2024, a particular flavor of optimism went around tech Twitter. It ran roughly: "GPT-4 is incredible. If only it were 10x cheaper, we'd solve everything. We'd be at AGI." Cost was the bottleneck. The intelligence had already arrived; it was just priced out of reach for the ambitious use cases.
The wish was granted, and then some. GPT-4-level quality fell from about $30 per million tokens at launch to under a dollar inside two years, a 200x collapse. It kept going. As of this week, June 2026, Zhipu's GLM-5.2 tops the Artificial Analysis Intelligence Index among open-weights models, priced around $1.4/$4.4 per million tokens and released under an MIT license, which means the marginal cost approaches zero if you self-host it. Frontier-grade intelligence, near-free, downloadable. The 2024 dream didn't just come true. It overshot by an order of magnitude.
Here is the part nobody dwells on. The person who made that wish is almost certainly not running a stack of near-free open-weights models for their important work today. They upgraded. They pay top-of-menu prices for Opus 4.8, or GPT-5.5, or whatever sits at the frontier this quarter. The cheap intelligence they begged for arrived, and they walked right past it to buy the expensive thing again.
That is the fallacy. Cheap models are not bad; they're excellent at what they're for. The fallacy is believing that getting today's frontier cheaply is the thing you actually want. It isn't. What you want is the frontier, and the frontier is never cheap, because "the frontier" is a moving target that always points at the most expensive thing in the room.
What happened to prices
Per-token inference costs have fallen faster than almost anyone predicted. Sam Altman, in early 2025, framed it as a law of its own: the cost to use a given level of AI drops about 10x every twelve months. Epoch AI sharpened the picture, finding that the rate of decline rose from roughly 50x per year before 2024 to around 200x after. DeepSeek dropped the floor again in early 2025 with frontier-class reasoning at roughly a 90% discount to incumbents. The trend has compounded since. GLM-5.2 in June is the clearest example yet: an open-weights model leading the open intelligence rankings at a couple of dollars per million tokens, free to download.
The cheap-model believer read the trend correctly, then. Where they went wrong was in what they thought it would buy them.
The floor drops faster than the ceiling
One pattern carries most of the argument: the price of "good enough" falls much faster than the price of "the best."
Take OpenAI's two ends over the GPT-4 era. The frontier, the best model on sale at any given moment, went from about $30 to roughly $2.50 per million tokens, call it 12x. The floor, the cheapest model that still cleared the old GPT-4 bar, dropped about 200x over a comparable stretch. GPT-4o mini matched or beat the original GPT-4 on most benchmarks at around 1/200th the cost. The ceiling barely moved in relative terms. The floor fell through the basement.
The clearest proof sits at the top of the market right now. Anthropic's flagship Opus pricing has held at $5 per million input tokens and $25 per million output since Opus 4.5, unchanged through 4.6, 4.7, and 4.8. Same sticker. The capability under that sticker climbed with every release: better coding, better reasoning, more reliable agentic behavior. The price of the frontier did not fall. The frontier got better at the same price.
That is how progress works at the top, and it's the reverse of the cheap-model fantasy. You don't get the same intelligence for less money. You get more intelligence for the same money, and yesterday's intelligence turns cheap as a byproduct, on its way to becoming a commodity the frontier no longer bothers with.
Altman conceded the point himself, almost in passing: pushing the boundary of AI won't get cheaper, but users will keep gaining access to more capable systems along the way. Pushing the boundary does not get cheaper. The frontier is structurally, permanently expensive. The cheapness only ever lands on last year's frontier.
"But GLM just shattered the frontier"
This is the objection that arrives the moment you publish something like this. On June 13, 2026, Zhipu released GLM-5.2: open weights, MIT license, a usable 1M-token context window, leading the open-model intelligence rankings, at a couple of dollars per million tokens or effectively free to self-host. Isn't that the cheap model finally catching the frontier?
It's the argument in disguise. The benchmark that crowned GLM-5.2 placed it on the Pareto frontier of intelligence versus cost, meaning the best you can get at that price. That is a different claim from the best you can get. The closed flagships still sit above it on raw capability. GLM-5.2 didn't reach the ceiling; it raised the floor again, harder and faster than before. The floor now moves fast enough that it can look like it's brushing the ceiling, but "best open model under $5" and "best model that exists" are separate sentences, and only the first one is true here.
There's a complication worth stating plainly, since it cuts the other way. Part of why GLM looked so frontier-grade that week is that the actual frontier briefly disappeared. The same day GLM-5.2 shipped, U.S. export controls suspended access to Anthropic's most capable Mythos-class models in some markets. When the ceiling gets pulled by policy, the next thing down starts to look like the ceiling. That's a story about regulation, not about cheap models catching up, and a reminder that "the frontier" is sometimes defined by what you're allowed to buy rather than what you can afford.
Either way the pattern holds. The cheap option got dramatically better and stayed cheap. The thing above it is still what people reach for when the work is serious, and now, in some places, it's the thing they aren't permitted to reach for at all.
Why you'll always want the sold-out model
Why don't rational buyers just pocket the savings and run last year's model at 1/200th the cost? On plenty of tasks they should, and I'll get to those. On their most valuable work they consistently don't, and the reason is economics, not status.
The first reason is that you pay for the outcome, not the tokens. When the work matters, production code that ships, a contract someone signs, an analysis a decision rides on, token cost is a rounding error next to the cost of being wrong. The metric that counts is price per correct outcome. A frontier model that resolves 69% of real engineering issues against 59% for a cheaper one, using fewer reasoning tokens to get there, can come out cheaper per resolved issue despite the higher sticker. The expensive model wins on the math that matters.
The second reason is that value at the top of the curve is super-linear. Altman's third observation was that the socioeconomic value of linearly increasing intelligence grows faster than linearly. A model that's 10% better is rarely worth only 10% more. At the hard edge of a task, that last increment of capability is often the gap between a usable answer and a useless one. The model that can actually do the thing commands a steep premium over the one that almost can, because almost is worth close to nothing when the answer has to be right.
The third reason is that we are cognitively greedy. When a new state of the art ships, demand snaps to it within days. Nobody weighing AI against the value of their own time asks for a slightly worse brain to save a fraction of a cent. People want the best brain available, especially when the other side of the scale is their own time, attention, or reputation. For high-stakes work, that instinct is correct.
The Jevons trap
There's a second reason the cheap-model dream curdled, and it should change how you budget.
When the price of something collapses, total spending on it usually rises. This is the Jevons Paradox, first described in 1865, when more efficient steam engines failed to cut Britain's coal use and instead sent it soaring, because cheap power made coal worth using for a thousand new things. Satya Nadella named the paradox during the DeepSeek panic of early 2025, when the market briefly assumed cheaper inference meant less demand for compute. The logic ran backwards, and Nvidia's stock recovered once the market caught up.
The AI figures are blunt. Menlo Ventures reports enterprise generative AI spending rising from $1.7B in 2023 to $11.5B in 2024 to $37B in 2025, a 22x jump in two years, while per-token prices fell more than 90% over the same period. The Silicon Data Token Expenditure Index roughly doubled in the back half of that run even as the price of a single token kept dropping. Prices fell about 90% and bills went up. Not a contradiction; the paradox doing exactly what it does.
So the original wish carried a buried false premise. "Make it 10x cheaper and our problems are solved" assumed the cheapness was the prize. It wasn't. Every time a tier of capability becomes affordable, a new tier of use cases becomes viable, and demand rushes in to fill the room the price drop opened up. You don't bank the savings. You find ten things worth doing that weren't worth doing yesterday, and the best of them want your best model.
When you should use the cheap model
None of this makes the cheap model a trap. The point is to match the model to the stakes. A cheap model is the right tool for a large and genuinely important class of work.
Reach for it when the task is high-volume and low-stakes: summarizing email, drafting routine replies, classification, tagging, autocomplete. When you run something a million times and no single output carries much weight, price per token is the right metric and the floor is where you want to be. Reach for it when "good enough" is a real, definable bar that the cheap model clears. If last year's frontier handles the job at 99% reliability, paying for this year's frontier buys you nothing, and you should take the savings. Reach for it when the cost structure is the product. If you're embedding inference into something with thin margins at scale, the cheap model is what lets the business exist.
The mistake runs in one direction only: reaching for the cheap model on the task where being right is the whole point. That's where the "I'll just use the cheaper one" instinct quietly burns more value than it saves.
The rule that holds
For your highest-value work, you'll want the most sold-out model. For everything else, the floor keeps falling, and you should ride it down.
A cheap model isn't the future arriving early. It's the past becoming affordable. Conflating the two is the fallacy.
The frontier is the most capable thing money can buy at a given moment, which makes it also the most expensive and the most in demand. It is always a little sold out. That isn't a supply glitch that cheaper hardware will fix; it's what the word means. The price of the best keeps refusing to fall while the best keeps getting better, and the buyer with something important to do keeps paying it.
So the next time someone tells you the real unlock is getting today's frontier model 10x cheaper, point out that they got their wish years ago, and ask which model they're actually running. It won't be the cheap one.