The Claude Devs account announced that starting September 14 they are “permanently raising” Claude Code’s standard weekly limit by 25%. Nice, except the current promotion gives users 50% more. So a permanent 25% raise means that two weeks from now users get about 17% less Claude Code than they get today. Anthropic said exactly that in the follow-up, to their credit. It still earned a Community Note.
The more interesting pricing trick is the other one, because it gets into what we are even buying when we buy AI.

ClaudeDevs@ClaudeDevs
Starting September 14, we're permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.
4:47 PM · Aug 29, 2026 · 6.84M Views
1.98K Replies · 988 Reposts · 18.6K Likes
Since Opus 4.7, Anthropic has been using a new tokenizer. Per their own documentation, the same input text produces roughly 30% more tokens than it did with the previous one. The exact difference depends on the workload. The sticker price is still per token. Same text, different number of tokens.
So how much is a token?
No one knows anymore.
We built almost the entire AI software economy on the assumption that a token is a currency. AI companies have real compute costs, so unlike traditional SaaS there is a direct marginal cost every time a customer uses the product. The frontier labs charged by tokens, so applications started charging by credits that translated, one way or another, back into tokens. Convenient chain: OpenAI charges me for tokens, my customer consumes tokens, I add margin in the middle.
Almost every part of that chain is now moving.
Earlier this week, Nebius showed GLM-5.3-Flash running at around 300 output tokens per second. Artificial Analysis currently measures Nebius at 301 tokens per second for the model, ahead of many other large open models being served today.
Good for Nebius, but it says something larger about the token economy. Since the early days of AI, we treated the model and the inference as one product. You called OpenAI, OpenAI ran the model, OpenAI decided what a million tokens cost. The application developer had little reason to think one layer below.
Open weights break that. If GLM is available to multiple inference providers, two companies can sell you the same model, generating the same tokens, with completely different economics underneath. One can optimize kernels better, batch differently, cache differently, use different GPUs, reserve capacity differently, and give you much more for much less.
The token didn’t change. The cost of producing it did.
We are going to care about this more than we do today, because inference is becoming a much larger part of the bill. Dell said this week that inference has already outpaced training demand and expects inference token volume to grow 87x by 2030, compared with 5x growth in training. The prediction can be wrong, and 2030 is far away in AI time, but the direction matters more than the exact 87x.
Training is something a small number of companies spend absurd amounts of money on. Inference is something every AI product pays for, every day, for every user.
Which raises a practical question. Would an AI company change its architecture because someone can give it cheaper inference?
If your model bill is $10,000 a month, no. You are not going to build an inference team because Nebius found a clever way to run an open model cheaper. OpenAI or Anthropic gives you an API; someone maintains the model, and you keep building your product.
Now take a company with agents or AI features costing millions of dollars a year. At some point, the CFO, the infrastructure team, or whoever owns that number will discover that, with some engineering investment, they can move part of the workload to an open model, a neocloud, reserved inference, or some combination, and materially change the product's margin.
They will make the move.
Maybe not all the way to owning GPUs. I doubt most application companies want that. But they will do some version of it, because once inference becomes a material line item, “we use this provider because this is what we started with” is not a financial strategy.

Wall St Engine@wallstengine
$DELL ON AI INFERENCE DEMAND: “Inference has passed training and is pure demand on our industry” Dell expects inference token demand to grow 87x to 3,600 quadrillion tokens by 2030, while training demand grows 5x “Enterprise agentic to be the single largest workload by 2028”

9:17 PM · Sep 1, 2026 · 175K Views
26 Replies · 123 Reposts · 839 Likes
This is also why gateways and routers become more important, not less, as companies get closer to inference. You don’t want every agent and product feature to know whether this request goes to Anthropic, GLM on Nebius, reserved capacity, or your own inference. You want that decision underneath the application.
And then metering gets harder.
At the top is the thing the customer thinks they bought: a support resolution, a generated video, a completed research task, a coding session, a playbook an agent ran successfully. This is where everyone likes to talk about outcome-based pricing.
Underneath that is the agent infrastructure: sessions, tool calls, caching, orchestration, retries, sometimes multiple models doing different parts of the job. Then a gateway or router, especially at any meaningful scale. Beneath that, inference, where someone is paying for GPUs.
These layers are connected, but not one-to-one.
You can cut inference cost in half without changing the customer’s outcome. You can improve caching and reduce the number of model calls. You can switch providers and make the same tokens cheaper. Or, as with Anthropic’s tokenizer, the same text is suddenly represented by more tokens without the customer getting 30% more product.
A pricing system that only sees tokens sees the easiest thing to count. I’m not sure it sees the thing you need to price.
fal pushed the same problem in a different direction. They released H3 Max, their post-trained version of MiniMax H3, which generates a five-second video in under three seconds. fal says that is around 35x the throughput of the official H3 endpoint and about 15x faster than models of comparable quality.
We can now generate video faster than the video itself.
People immediately did the obvious internet thing with it. Rehan Sheikh from fal wired it into an infinite generated TV stream, so while you watch one piece of AI slop, the next one is already being generated. I found it weirdly fascinating.
There is a serious pricing point inside this. When a five-second video takes a minute to generate and someone makes it take 30 seconds, you have a faster video model. When it takes less than five seconds, you can build a live product. Speed stopped being an infrastructure metric and changed what the application can do.

@levelsio@levelsio
Okay I built it! 🍰 Infinite Slop levels.io/infinite-slop An infinite and interactive AI generated live stream of slop that goes on forever and ever Anything that you write in the chat is generated next and AI will try to connect it to the previous video so there's an actual…


@levelsio @levelsio
Today is a very historical moment for AI video generation You can now generate AI video faster than you can watch it Before it'd take let's say 2-5 minutes to generate 15 seconds of video @fal made a post-trained Minimax H3 variant called Max which is 50x faster than the
5:34 PM · Aug 29, 2026 · 2.43M Views
862 Replies · 553 Reposts · 6.88K Likes
That reminded me of a conversation this week with the CEO of an AI company. We were talking about authorization and governance inside their agents, and I was worried about adding latency to a complex authorization decision. His answer: I don’t care. If the user asks an agent to do a complicated task and already expects 20 or 30 seconds of work, another second for authorization changes nothing. He is right, at least for his product.
So latency itself is not the hard thing to price. The hard thing is knowing when latency changes the product.
For one agent, an extra second is worth zero. For fal, crossing the point where generation beats playback creates a new class of application. The same goes for throughput, concurrency, priority, availability and probably other units we haven’t cared much about yet.
We know how to do this in SaaS. Quality of service has always been part of pricing. You pay more for uptime, more capacity, faster queues, larger concurrency, better support, priority processing. Cloud infrastructure is full of these distinctions.
AI collapsed most of that into tokens, because tokens had such a clean connection to the model provider’s cost.
But if one inference provider gives you 300 tokens per second and another gives you 100, and for one workload the difference doesn’t matter while for another it enables the entire product, then “one million tokens” tells me very little about the value being delivered.
The real unit of AI economics might be tokens in some cases. It might be speed in others. Guaranteed throughput, successful tasks, seconds of video, concurrent agents, or some combination.
Which brings me back to the annoying part: your metering system probably needs to understand the whole stack. You can’t look only at what the customer consumed, and you can’t look only at what the model provider billed.
The biggest confirmation this week that inference is moving closer to the application layer came today, when NVIDIA announced it is acquiring Hugging Face for $12.93 billion. Hugging Face has more than 18 million developers, more than 3 million models, and more than 200,000 companies using the platform to discover, evaluate, customize and deploy AI. NVIDIA says Hugging Face will remain open, multi-cloud and multi-accelerator, and that NVIDIA compute will not be required.
Calling Hugging Face an inference company would be wrong. It is the place where a huge part of the open-model world discovers models, distributes them, evaluates them, fine-tunes them and figures out how to run them.
NVIDIA already owns a very important part of the layer underneath.

Jensen Huang@JensenHuang
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you…
12:02 PM · Sep 3, 2026 · 824K Views
723 Replies · 1.4K Reposts · 10.7K Likes
So consider the position. On one hand, the company selling a huge share of the compute used to create tokens. On the other, that same company now owns the place where millions of developers choose which open weights they want to turn into tokens.
This doesn’t mean NVIDIA can decide what a token costs. Hugging Face staying open to other clouds, accelerators and inference providers matters here. But NVIDIA is now much closer to both sides of the exchange. They sell the commodity, and they own one of the most important places where developers decide what to run on top of it.
For companies running open models, this is a very different world from “Anthropic says a million tokens cost $X.”
The model can move. The inference provider can move. The GPU can move. The tokenizer can move. The amount of caching can move. The required quality of service can move.
The token looks less like the currency and more like one measurement somewhere in the transaction.
Two other things from this week make it more interesting.
The first is UCP, the Universal Commerce Protocol, a protocol for commerce between applications, businesses and agents. It gives agents standardized ways to discover commerce capabilities and perform checkout, with bindings for REST, MCP and A2A. The specification and implementation are open on GitHub.
Most agentic payments won’t be agents paying for compute. But a meaningful amount of machine-to-machine commerce will eventually be software buying software and infrastructure on behalf of a user or another system. When the buyer is an agent, switching suppliers gets much easier. An agent can weigh capability, price, latency, privacy requirements and availability, then choose a provider with no emotional attachment to the logo on the API endpoint.
Then there is Ornn, which anyone working seriously on compute economics should look at. They call themselves “the foundation of the compute market,” and they are trying to treat GPU compute like a real commodity market. Their OCPI tracks traded spot prices for hardware such as H100s, H200s and B200s, and the larger idea is to make compute something companies can price, finance and eventually hedge instead of buying through an opaque collection of contracts
The combination of these two directions is what interests me. UCP works from the application downward: software can discover something it needs and transact for it. Ornn works from the infrastructure upward: compute should have a transparent market price.
Somewhere between them sits the token, which we use as if it is the natural currency of AI.
Maybe it was, for the short period when the frontier lab, the model, the inference infrastructure and the price were one package. Once those pieces separate, the token becomes a much weaker abstraction for money.
I’ll still count tokens. Everyone should. They are useful for understanding model usage and they are still how a lot of providers bill us. I just wouldn’t confuse counting tokens with understanding the economics of an AI product.

