Agentic coding is where large language models found product-market fit: agents burn vastly more tokens than chat, and it is a useful tool for many professionals including me.
A year ago $20 a month was plenty; today $100 is the entry bar for serious work. Marty Kausas, CEO of Pylon, admitted: “I accidentally spent $4,000 in 3 days in Claude Code.” Uber rolled out Claude Code in December 2025 and burned its entire 2026 AI coding budget by April. The Information reports multiple enterprises facing bills 2-3x higher.
Seats with usage included vs pay per token
The same work, priced per seat (Max, Team) and per token (API, Enterprise):
| Report | Setup | Paid | API-equivalent | Multiple |
|---|---|---|---|---|
| Author, June 2026 | Claude Max 5x | $100 | $2,092 | 21x |
| Author, July 2026 | Claude Max 20x | $200 | $2,986 | 15x |
| Simon Willison, May 2026 | Claude Max 5x | $100 | ~$1,200 | 12x |
| SemiAnalysis test, June 2026 | Claude Max 20x | $200 | ~$8,000 | 40x |
| Pylon, whole org, June 2026 | Claude Team vs Enterprise | $400K/yr | ~$1.4M/yr projected | 3.5x |
These are list-price math, not real bills. Self-reported numbers on Hacker News land in the same 12-50x band: from $1,850 a month at half the limits of a $100 Max 5x, up to “$15k in the past 30 days” on roughly $300 of subscriptions.
Organizations where many people use Claude irregularly, or for non-agentic work, see more favorable math. But among the companies I have talked to, everyone who moved from seats to per-token Enterprise saw the bill at least double, and most reported roughly 3x.
The realistic price for agentic coding
| You buy | What it costs (August 2026) |
|---|---|
| Seats: Max, Team | $100-200/user/month, usage included within limits; Team capped at 150 seats |
| Tokens: API, Enterprise | Opus 5: $5/$25 per million tokens Fable 5: $10/$50 per million tokens Cache writes at 1.25-2x input Cache reads at 0.1x input Enterprise adds $20/seat |
Full details on Anthropic’s pricing page. The same models can also be bought per token through the major clouds (AWS Bedrock, Google Vertex AI, Microsoft Foundry). The non-obvious part of per-token billing is the cache: agents re-read their whole context on every step, so long sessions are mostly prompt-cache traffic. Across all my sessions, 82% of the API-equivalent cost was cache related. Though this is good for usage, it makes bill harder to analyze.
How companies cap AI spend
A per-engineer cap with an override path is the current norm among heavy adopters:
| Company | Monthly cap per engineer | Source |
|---|---|---|
| Uber | $1,500 per AI coding tool, exceedable with permission | Bloomberg, June 2026 |
| Workday | ~$2,000 | SemiAnalysis, June 2026 |
| Stripe | ~$2,000 | SemiAnalysis, June 2026 |
| Atlassian | $500-2,000 “AI wallets”, tiered by role | The Next Web, July 2026 |
| CloudZero | $5,000, sized so it never binds | CloudZero, May 2026 |
| Shopify | No cap: an alert at $250/day, investigated rather than restricted | Bessemer, April 2026 |
These are the heavy adopters, not the average: in the Pragmatic Engineer survey, the typical company-funded plan is $100-200 per engineer per month, and Gartner found nearly a quarter of tech leaders spending $200-500 per developer per month on AI coding tokens, with only about 6% above $2,000.
Why Anthropic leaves money on the table
My understanding: the seat plans are a subsidy and market segmentation. Cheap seats let people learn agentic coding, get good at it, and shape both the product and their own preferences, so that adoption later happens at much bigger scale. Depending on what you count, subscriptions bring in only 5-15% of Anthropic’s revenue: SemiAnalysis estimates consumer subscriptions alone at ~5%, while Sacra puts all subscription plans combined at 10-15%. The metered side, dominated by enterprises, is where the money is made.
Those multi-billion-dollar data centers full of chips are not cheap: data center capex surged 57% in 2025 and is forecast to top $1 trillion in 2026. Anthropic already tried to meter Fable 5 for subscribers before settling on including it in Max plans at up to half the weekly limits, and it will likely try again when competition allows.
What Claude Enterprise actually is
Claude Enterprise today is a roughly $20 per seat license that includes no usage at all. Every token bills at standard API rates on top: self-serve customers prepay into a shared credit pool, sales-assisted customers get monthly invoices after usage. There is no published volume discount, though there are rumors of discounts to smooth the price-hike transition.
When Claude Enterprise is a must
Seat pricing stops working in three places:
- The 150-seat limit of Claude Team. Past it you must move to Enterprise, where Anthropic earns the most.
- Enterprise-only features. SCIM provisioning, audit logs, and compliance and analytics APIs for exporting per-user usage live only on the Enterprise tier. Anthropic also has a habit of parking capabilities there: a longer context window was an Enterprise exclusive when the tier launched, and today Mythos, the ungated sibling of Fable for approved use cases such as cybersecurity, goes only to approved enterprise customers.
- Gotchas of consumer plans at work. Individual Max accounts mean individual invoices, which procurement hates. Consumer plans lack commercial terms and a data processing agreement, the model can be trained on your data if you click the consent prompt the wrong way, and in the EEA and Switzerland the consumer terms even carry a non-commercial-use clause, mostly a liability limitation (discussion).
How to manage Claude Code costs
Give people Claude Max and think of it as training, like a conference ticket. Many early agentic projects underdeliver against the executive scope. That is fine: using the tool is the only way to learn it. Fund it as a perk or reimburse the subscription, side projects included: you are learning on the subsidized tier instead of at API prices. Cancel the seats nobody uses.
Grow organically from Max to Team, and move to Enterprise only once you outgrow 150 seats. If your company has multiple divisions, buying multiple Team workspaces is also a great path; centralization is an anti-pattern.
When someone hits the limits and cannot work, the first instinct should be another subscription, not API tokens for the overflow. API bundles can offer discounts of up to 30%, but that does not close the gap with subscription pricing. Plenty of people run more than one: a Max for experimentation and a Team seat for commercial work.
Some companies use a second vendor to stay under 150 seats: core engineers get Claude Code, everyone else uses OpenAI Codex or another provider. Both groups keep seat pricing, and the multi-vendor setup is negotiation leverage besides.
Before moving to Enterprise, get good at cost monitoring. Once you pay per token, per-seat budgeting is crude. Agentic spend varies wildly, and your biggest spenders are often the people using the tool the most. At the same time, we are all still bad at measuring the impact. A weekly budget per person, actually measured, plus a lightweight process to raise limits for the people who deserve it, beats any flat cap.
Claude Code budgeting antipatterns
If you hand people the latest Fable model with ultracode multi-agent tooling and a small allowance, they will burn the weekly budget in a few hours. That is a terrible first experience. Better to default to a slightly weaker model, or a lower reasoning effort, that people can use all week than something that regularly cuts them off mid-task. Claude Pro and small Team plans are the official-packaging version of the same mistake: at those limits agentic coding is barely usable, and Claude Code likely stays on Pro mostly for PR reasons (Anthropic tried removing it in April 2026 and reversed within a day after backlash).
If money is short, restricting Fable and defaulting to Opus helps, but there is a limit to how much you can save inside Anthropic’s price list: Anthropic in 2026 is Apple, not cheap, but many like it the most. There is a whole competitive field beyond it, including cheap Chinese models (DeepSeek, Kimi, Qwen, GLM) hosted by Western companies and alternative harnesses (Pi, OpenCode), but that deserves a blog post of its own.
The cheapest model has uses: permission checks and simple mechanical tasks, where it is unbeatable per dollar. For long-horizon agentic coding it is not great, and people restricted to it get a poor experience of the whole technology.
The opposite failure mode is letting everybody burn the whole budget in a short time with no visibility, and then discovering the money is gone without knowing what it went to. That is how Uber blew its 2026 AI coding budget in four months, with its CTO telling The Information: “I’m back to the drawing board, because the budget I thought I would need is blown away already.” The typical answer is a per-engineer weekly/monthly cap.
Plans change, promos lapse, and models leapfrog each other within months. Patterns that work elsewhere, like centralizing procurement to save money, can backfire badly here.
Token economics matters
But the economics matter now. You can spend more on tokens than on the engineers driving them, and spend can scale to almost any number if nobody is watching. The tooling for visibility still lags what enterprises need.
If you are wrestling with an AI bill, I would love to hear your story: contact me through e-mail or a call.