Historically, land was the most valuable asset on the globe. Political conflicts were waged to gain control over land, and when an excessive amount of land was accumulated in a few hands, societies divided into aristocrats and commoners. During the industrial age, land was replaced by machines and factories. The focus of the conflict then shifted from land to the means of production, and society divided into capitalists and proletarians.
The main question that follows is what is the most important asset in the AI age?
A useful feature of this historical pattern is that it maintains the structure and order even through all the changes. In every era, the most valuable productive asset is a claim on other people’s necessity. The owner of the means that people cannot live without gains from everyone who is affected by it.
Who owns the means that are used by everyone in this era, and who owns the tollbooth?
Of course, compute is the most likely candidate to compare with land. And it deserves attention because, to a great extent, the analogy with land is valid.
The numbers are land-like in scale. Goldman Sachs’ baseline model suggests $765 billion in annual AI infrastructure spending for 2026. By 2031, they expect annual spending to grow to $1.6 trillion. CreditSights estimates the top five hyperscalers would spend $602 billion on AI computing by 2026, over $400 billion more than in 2025, with about three-quarters of that spending going to AI infrastructure. Broadcom, Inc. has stated publicly that the primary supply chain constraint for them is the capacity of Taiwan Semiconductor Manufacturing Company (TSMC) and that they expect the supply to improve in 2027.
With that constraint, compute resembles land. It can be extracted or leased rather than owned, and compute owners earn rent regardless of what gets built on top. Passive income is the name of the game.
It’s worth noting that land was fixed and compute is manufactured, and manufactured goods respond to price. TSMC expects to ramp up their 2026 capex by 40%, setting their target at $52-$56 billion. The CHIPS Act alone has facilitated over $600 billion in private commitments.
Fabs get built, nodes ramp, and the queue shortens.
While the cost of compute is plummeting, Epoch AI found that for certain tasks the cost to reach a milestone capability drops between 9x and 900x per year. Costs to achieve GPT-4 level performance on PhD-level science questions, for example, drop by roughly 40x every year. Current rates do the talking. DeepSeek V4 Flash lists for around $0.14/$0.28 per million tokens. Closed frontier models cost around $5/$25.
Durable toll booths do not behave like that. Commodity inputs in the process of commoditizing certainly do. Current scarcity is real and has an expiration date.
Then it must be the models themselves, you might say. The labs that turn compute into intelligence.
By all accounts, though, it isn’t the companies with the largest compute, cash, or user bases that control the frontier. The incumbents with access to unlimited resources and planet-scale databases should have easily won. Mostly they didn’t. The small labs with none of these advantages did.
Whatever the advantage ended up being, they were not purchased by these inputs.
What those labs designed and continually built is the loop: a circuit that converts deployment into a better model and a better model into more deployment.
Here is what i mean when i say a loop, this is what it consists of. Four things must be connected end to end for the circuit to close.
A deployment surface, where the model acts in the world.
A signal, which has to be definable. That is, it has to establish whether an action was deemed correct versus if someone liked reading it.
A signal training pathway, which is a route the signal takes to reach the weights on a fixed schedule.
A win, which happens when the improved model creates more deployment and more signal.
With any one of the four broken, what you have is a good product with analytics.
Customer support is the best example. An answer is given, and the system knows if the answer was correct, and that feedback trains the next version, and the next version resolves more tickets, and those tickets produce more feedback. The loop keeps compounding, and someone always needs to be putting money into it. Whatever they put money into is what the system optimizes.
The technical name for the well-behaved version of this is reinforcement learning with verifiable rewards, RLVR, which sits behind most reasoning models shipping today. Its defining property is the one that matters here: it works where ground truth exists and degrades into noise where it doesn’t. One widely circulated result had a Qwen math model gaining 21.4% on MATH-500 from random rewards against 29.1% from real ones, which should worry anyone who believes a thumbs-up button counts as a training signal.
Worth holding onto. I’ll come back to it later.
The loop sorts people into two broad groups.
The first group is made up of individuals whose goals intelligence serves. The second group is made up of the behavior, attention, and output variables that are inside someone else’s optimization.
I like to think of this as the aristocracy and commoners, or the capitalists and proletarians, but you could also call it the owners and feeders.
Owners keep the loop while feeders do the filling. Any adjustment a feeder attempts is recorded as a balance on someone else’s books. Owners of the circuit a labeled gradient flows through accumulate the value of every suggestion accepted, draft corrected, and ticket closed.
The division is included in the terms of service, which most people overlook. The walling off of the division in enterprise contracts has largely been accomplished: OpenAI does not train on API or enterprise data by default and has not done so since March 2023; Anthropic provides zero-data-retention arrangements to commercial clients. Consumer contracts have been going in the opposite direction. From September 2025, Anthropic’s consumer chats and coding sessions will provide training unless you opt out, and retention will be increased from 30 days to 5 years if you opt in.
Read the two policies, and you will see the class division in the contracts. If you can negotiate your terms, you are not a feeder. If you click Accept, you are.
“But this is not new” is a fair response. Since the beginning of industry, transferring skill into capital has been part of the process. The designs for factories were a replication of what was already known to be possible.
Actually, it’s the speed and precision that have evolved. The factory learned a craftsman’s skill over many years by purposefully studying and replicating. Someone always had to watch you, analyze your process, and construct the next step around your work. Here it’s the loop that auto-learns at a very high speed with each interaction. Every correction becomes training. No one has to watch you. You train the system by simply doing your job well, and the data labels itself. In non-industrial white-collar work, that is the real alpha that is produced by hiring the right people, getting the right deals, understanding the market, and much more that enterprises outside of physical industries have built over hundreds of years.
The industrial version required intent and a dedicated department. This version requires a webhook.
“But I’m getting paid in productivity” - That’s a great argument.
People say: yes, I am feeding someone else’s training loop, but I receive tremendous value and really, it’s a good trade.
It could be. The real test, however, is whether the productivity gains occur for you faster than the increase takes place for the capability against you.
Your gains are limited and local, task by task and ticket by ticket. They are generally reset by changing jobs or tools. For the loop, this is not the case; it always retains the gains from each correction across all feeders, and they are compounded.
So, in the end, the trade is to support the thing that will eventually make your work obsolete, and move the alpha away from productive enterprises toward the organization that captures or can pay for the signal.
In practice, what you call the ‘democratization of consumption’ is much more narrow. Here, the phrase applies to the ability to consume intelligence. And few have the ability to create it.
And right on cue comes the counterexample. Open source. The designs are public. Anyone can access them. And Chinese labs are shipping cutting-edge models with a permissive license. Alibaba announced Qwen had 700 million Hugging Face downloads by January 2026. A Linux Foundation and Meta study reported that 89% of the organizations that use AI at some point use open-source models. Is the counterexample closed?
Not necessarily. And the most interesting number in this whole article is right next to that 89%.
Menlo Ventures researches the percentage of real enterprise LLM workloads utilizing open models. It was 19% in 2024. In the first half of 2025, it fell to 13%. At the end of 2025 it was 11% in a market that grew from $11.5B to $37B with three companies capturing 88% of the market for enterprise LLM API use. You might argue that it has grown since then, especially in 2026, but that’s not the case for large-scale or non-technical sectors. Most of the growth is experimental right now, with tech-first companies and startups adopting more rapidly due to pricing pressure and use case maturity.
Open-source models are free, nice, leading-edge on some benchmarks, and they cost between 10 and 100 times less per token. Almost nine in ten organizations have come across them. And in the last two years, the percentage of real enterprise workloads has been falling.
This is not a capability problem. Then what was it?
Treating two different assets as if they were one creates the confusion.
We need to think about our model as one asset, while the infrastructure around it is another. You can obtain the first asset with Open Weights, but not the second. As it is common practice to post high-impact models for free with no context along with them, you might receive, for free, a model representing hundreds of millions of dollars worth of computing time. You do not receive the infrastructure necessary for evaluation and deployment, nor the cadence with which they are performed, that would contribute to the asset’s continuously increasing value.
What would a world of true open weights be like? Two categories.
The first would be at the frontier, consisting of labs that operate at an extremely large scale. Running underneath them would be a large population of players that conduct self-sovereign loops by deploying an open source or containerized frontier model on their own data, obtaining their own feedback, optimizing the model within the domain, and providing no feedback to any external model providers. This is a real form of democratic production, and it is largely absent.
True open weights are about self-sovereignty. You run the model, and your data are yours. However, if weights are run outside the maker’s control or their built loop, two things happen. The usage signal does not flow back to the maker, which is the intended consequence. The company running those weights also cannot benefit from the usage, since they have no way to capture the signal and no pipeline to regularly use it.
To prevent the signal from being diffused. It gets destroyed. No one captures it.
Current usage of open models basically means having a model without a loop. And that, rather than benchmark scores, is why the 11% continues to decline. Enterprises are selecting the model that improves on its own over the model that they own but have frozen. They want serviceability.
Many people argue with me on this that fine-tuning already solves this issue. It’s cheap, LoRA made it accessible, and it’s fast now. When NVIDIA shipped Nemotron 3 Ultra with open weights, Harvey reportedly post-trained it for legal work in about 24 hours.
Nobody disputes any of that. But a fine-tune is a one-shot injection of what you already knew. It’s a snapshot rather than a subscription, and it gives you no accumulation. Also, the funny part is that a lot of enterprise data and enterprise work, especially in non-technical industries, has never been captured. You also don’t have a lot of enterprises that have good enough pairs to even find you, so the journey begins at the capture.
Let’s break the counterfactual. Your fine-tuned model is used by hundreds of users over thousands of interactions. After getting hundreds of instructions, some of the instructions go bad over time. Your business systems know that invoices were reconciled, tickets were closed, models stood review. All the verified feedback is just lost because nothing kept it in infrastructure or captured it carefully, and it just died.
The finetune moves you once. The loop moves you every cycle, compounding on a signal your own deployment generated.
So the conclusion isn’t that open source needs better models. It already has better models than it can currently use. What open source needs is an infrastructure.
Placement matters. This is why most variants fail and disappear.
The harness captures an enterprise’s own verified signal and feeds that context to any intelligence, be it frontier or open weights. The pitch is not “contribute your data to make open or frontier models better for everyone.” The pitch is own your loop. The system’s design is such that serving multiple enterprises from a single harness is what justifies the system’s construction economically, yet the signal remains unpooled.
Placement is the most important design choice, and most fail it.
The harness must be placed at the verification surface instead of the chat surface.
If thumbs-up are captured from a chat interface, a preference signal is generated that is weak, noisy, mostly about tone, and is very much open to the spurious-reward failure mode described above. This would teach the model that being agreeable is correct, but in many cases, an agreeable model is an actively dangerous one.
This is the part our own research made concrete for me. In Where Does Agent Reliability Come From? we used a frontier model in an execute–observe–compare–correct loop with small specialist checkers. SpreadsheetBench Verified (n=400, p<0.001) improved from 80.25% to 91.25%, BullshitBench v2 improved by 7 to 10 points, and GAIA improved from approximately 60% to 75.2%, with verification costing 2–10% of the frontier-model cost. The loop with the trained specialists was the source of the improvement. Returning the checking to the model that created the artifact eliminates almost all of the improvement.
The important reason here isn’t related to reliability, which was the focus of the study at that time. In verification loops, every check adds a label to the judgment about the correctness of the output. If you already run one, you are already creating the exact training data that the labs are spending billions to create, and you are discarding it.
Most enterprises are sitting on these surfaces already:
The CI pipeline that says the code passed.
The ticketing system that says the issue was resolved and stayed resolved.
The ERP that says the forecast matched actuals.
The reconciliation that closed, the model that survived IC review, the valuation that held at close.
These systems produce verified labels daily for every business, in quantities that no labeling vendor could possibly compete with. Nearly none of this is utilized for training.
That is what the industry is spending on now. Artificial RL environments have essentially become a funding category of their own. It’s been reported that Anthropic has discussed spending over $1 billion on RL environments within one year. Prime Intellect’s Environment Hub crossed the 1,000 community environments milestone in April 2026. Surge went over $1 billion in revenue, while Scale was valued around $29 billion.
It’s a fact that billions are invested to replicate the ground truth that your ERP makes for free. This funding asymmetry is the opportunity.
Capture still matters, especially in a world with million-token context windows. Packing your data into a prompt is a pay-per-use capability. There is no saving for prompt use. The only ways to save are to do specific context augmentation or train on it. Foundation Capital believes that 2026 will be the year that decision traces will be the moat, and small customized models will outperform frontier models in the tasks that have the most value in enterprise infrastructure.
Capture is related to the tenant. Training can’t be, and we will not keep this project idea alive by pretending otherwise.
Retraining a model is more of a specialized task than a script. You need to capture the right signal and balance the mixture. You risk losing what the model already knows, and catastrophic forgetting will occur. It is well documented across models of 1B to 7B parameters. It will become worse as the model size increases. After every retraining, you need to evaluate whether you improved the model or if you just changed it.
Training models well is not something the majority of mid-sized companies can justify building in-house, as they run it maybe only once a quarter. The economics do not work, and this, not model quality, is the real reason the sovereign tier does not exist.
If we are going to see a lot of sovereign deployments of open models, things need to change, and we need to treat training as a utility. That would mean we capture a lot of signals in isolation and provide tenants with updated models on a regular basis.
As we have implemented a shared system, we assume trust is no longer an issue. We need strict requirements and structural guarantees, not contracts. We need architecture, not promises.
It sounds impossible until you realize the semiconductor industry solved a harder problem 40 years ago. Fabless chip companies hand TSMC their most valuable trade secrets (IP), years of research and development (R&D), and the most valuable part of their business. TSMC manufactures it and gives it back. They run over 500 customers through their facilities with no leakage, and have done so since 1987 with one structural commitment: they never compete with their customers. The rule is pure-play; they completely avoid making their own chips.
That one rule created the entire fabless industry. There is no NVIDIA, no Apple Silicon, no Qualcomm without a manufacturer that credibly doesn’t want to become them.
Likewise, a training utility requires the same sort of rule. Not a model company that does training as a side business, but a pure-play trainer whose entire franchise is built on never using their signal for anyone else, including themselves.
Let’s go back to 11 percent. It breaks down neatly now.
Enterprises don’t opt for closed models because of capability limits. Open models have crossed the threshold for a long time. They go for closed models because they are managed. The model gets better; someone else does the retraining, evaluations, and migrations, and you wake up to a better system having done nothing.
This is why “free weights, 20x cheaper per token, full sovereignty” loses. You are given a static asset in a market where the alternative, which is more liquid and self-appreciating, is more expensive, but not frozen. Enterprises go for more self-appreciating assets, and they are actually making the right decision.
Do the opposite and put a shared but tenant-isolated infrastructure behind open models or inference providers, and the whole calculation inverts. Your open deployment gets better, and it does so based on your own signal and in your own domain, and at your own pace. It does it without your data leaving your control, and it improves along the axis you actually care about, which is your work rather than the average of the internet’s.
This is what would make open models a real choice on serviceability instead of cost and control. Both have been true for two years, and the share still fell.
The real question of this era is not who has the best model. Models diffuse. Weights leak, get published, get replicated, and within months the best open and closed models are separated by a few months.
The main question is who owns the loop that the model sits inside. I believe for most of the economy, the answer is going to be “someone else.”
That is currently the case. Open Weights didn’t fail. We shipped the artifact and left the circuit out of the box.
More on the structure of loop infrastructure in the next post: what a tenant-isolated capture harness looks like in production, where it goes in an existing stack, and what a pure-play training utility would have to guarantee.
Until next time,
- Arunabh Dastidar







