An introduction from our CTO
The Māori language Te Reo has no commercial market to speak of, but a broadcaster in New Zealand's far north has been building speech models for it anyway — and all under a license that keeps the recordings with the communities that originally gave them. Half a world away, in a cellular dead zone in East Africa, farmers point a phone at a cassava leaf and get a diagnosis from a model small enough to live on the handset. Neither project needed permission. And neither could have been rented from a frontier model.
The largest companies on earth are now thinking this way. This month, on the Artificial Analysis Intelligence Index, the strongest closed model scored 61. The strongest open model? 57. The open model came in fourth overall, ahead of models from three of the biggest closed labs. At the end of 2025, about a third of all tokens on OpenRouter were being routed to open-weight models; now the seven highest-volume models on the platform all ship open weights.
In a surprising recent turn, corporate heavyweights are leaning in an encouragingly unexpected direction: On July 24, Mozilla signed an open letter alongside NVIDIA, Microsoft, Meta, IBM, Dell, Mistral, Hugging Face, the Linux Foundation, and Andreessen Horowitz that argued that open weights are too important to be ignored.
We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community made sure it never could. We bet on open the first time. Open won. Together, we can do it again.
Raffi Krikorian · Chief Technology Officer, Mozilla
01The current state of open models
The model layer has commoditized.
Value accrues to the harness above it.
Open weights are where the work happens. A majority of production tokens now route through them, and the seven highest-volume models on OpenRouter are all open weight. Closed models still lead at the frontier, on reasoning and multimodality. However, most production workloads run well below that ceiling. Commodity inputs surrender pricing power. Value moves up to the agentic harness, the layer above the model.
Terminology note
Open model here means weights you can download, run and modify on hardware you control. That category splits. Open weights means the parameters ship under a permissive license with no training code and no data documentation, which describes most of what this report measures. Open source AI, as OSI defines it, additionally requires the training code and enough information about the data to rebuild the system.
The frontier's top three are closed. The fourth is open.
Artificial Analysis Intelligence Index v4.1, a nine-eval composite that includes Terminal-Bench 2.1, Humanity's Last Exam and GPQA Diamond. Kimi K3 ranks fourth, four points off the closed frontier.
Open weightsClosed
4 ptsfrom the best open model (Kimi K3, 57) to the closed frontier (Claude Opus 5, 61)
3 of 11top-tier models shipping open weights
Source: Artificial Analysis Intelligence Index v4.1, July 2026, top 11 of 586 models shown. Open = downloadable weights.
3.6 points off the top, at a third of the price
The same index against weighted-average price per 1M input tokens, for the top ten. Kimi K3 is the only open-weight model in the leading cluster, fourth overall and 3.6 points off the top, at about a third of the price.
Source: Artificial Analysis Intelligence Index v4.1 via OpenRouter, July 2026. Price per 1M input plotted where list price is disclosed.
Open trails the frontier by 6 ECI points, about one release cycle
Epoch Capabilities Index by release date. The best open model, Kimi K3 at 156, against the closed frontier, GPT-5.6 Sol at 162.
Source: Epoch AI, Epoch Capabilities Index (CC-BY), July 2026. Open = downloadable weights. The gap is the point difference at the frontier, and the confidence intervals overlap.
A jagged frontier: parity, contested, a closed edge
For most workloads open models already clear the bar. The closed frontier earns its premium in a narrow band, namely deep reasoning, long-context reliability and professional-knowledge polish. Match the model to the job and you need the frontier for less than you think.
Open leads or parity
Frontend coding
K3 debuted first on LMArena's Frontend Code Arena at 1679 Elo, leading in six of seven frontend domains, plus coding, instruction-following and general knowledge.
Frontend Code Arena measures building website and UI code, scored by blind developer votes.
Contested
Agentic terminal work
K3 scores 88.3 against Sol's 88.8 on Terminal-Bench 2.1. It wins Program Bench, SpreadsheetBench 2 and BrowseComp, and loses FrontierSWE at 81.2 against Fable 5's 86.6.
Terminal-Bench, Program Bench and BrowseComp measure agent work, running terminal tasks, writing programs and researching the web. FrontierSWE measures resolving real software-engineering tickets.
Closed edge
Professional knowledge work
Fable 5 leads K3 by 92 Elo on GDPval-AA v2, the largest Elo separation among the shared benchmarks, alongside long-context fidelity and conversational polish, which Moonshot concedes still trails.
GDPval-AA v2 measures expert-graded professional knowledge work, where long-context reliability and polish appear.
Scores use different scales and come mostly from vendor-run tests, so treat them as directional. LMArena uses blind human voting. Sources: LMArena, Artificial Analysis, Moonshot K3 blog.
Inference fell 50× in 36 months
Cheapest model at GPT-4-class performance, blended API list price, log scale, against prior platform-shift cost curves at their historical rates. Over the same 36 months the dotcom bandwidth curve delivers 2.6× and the PC compute curve 3.4×. The frontier price fell 112× from GPT-4's $45 launch.
Sources: a16z "LLMflation" methodology, Epoch AI LLM inference price trends, Artificial Analysis, Moonshot, MTS/Substack.
Open weights win the tokens
The share of tokens routed on OpenRouter through open-weight models grew from a negligible base to a third by late 2025 to a majority by mid-2026.
Source: OpenRouter 100T-token study (Nov 2024–Nov 2025) and live leaderboard, with intermediate points interpolated. By request count, closed US providers still lead. The open lead is a token-volume lead, concentrated in coding and agentic workloads.
Open weights dominate token usage
Each model's share of the top-20 routed token volume on OpenRouter, 1–27 July 2026. The top seven by volume are all open weight. Anthropic's closed Claude models are the next entrants.
Open weightsClosed
Top 20 tokens, by licence
Open, ranks 1–10 · 72.4% 8.7% 18.9%
72.4% Open, ranks 1–10
8.7% Closed, ranks 1–10
18.9% Ranks 11–20
K3 note: the API opened 16 July and the weights on 27 July, so K3 does not appear in July's routed-token top 20. New subscriptions were temporarily paused on 20 July as demand neared capacity. These ten account for 81.1% of top-20 volume, leaving 18.9% across the remaining ten. Shares are of the top 20 only, not of all OpenRouter traffic. Source: OpenRouter LLM Leaderboard, July 2026 (this month), routed traffic only.
The open frontier and the deployable open frontier split
Minimum serving footprint per model, measured in 8-GPU nodes. Bar length shows nodes to run, and memory is the figure that produces it. Frontier capability is downloadable. Node count decides who can run it.
Node counts follow Moonshot's guidance of 64+ accelerators for K3. A single 8-GPU node barely holds K3's weights, so eight nodes reflects serving it rather than only loading it. Inkling's ≥600 GB straddles the one-node line depending on GPU generation. K3 open weights released 27 July 2026. Sources: Moonshot deployment guidance, TML model card, Hugging Face community, Northflank.
Open ships easy.
Open deploys hard.
Data from the Mozilla / SlashData 2026 developer survey. Open models lead in adoption. 79% of developers adding AI functionality use them, against 71% for closed, and the two are largely complementary, with half of developers using both. But production is where teams stall. Only 53% of open-model teams reach production versus 63% for closed. The gap traces to operational tooling and trust.
Open models lead in adoption, and mostly coexist with closed
Share of developers adding AI functionality to their applications who currently use each model type, and how the two overlap.
How they combine
29%OS only
50%Both
21%CS only
Source: Mozilla / SlashData 2026 developer survey. Most teams treat open and closed as complements, with 50% running both, 29% open only and 21% closed only.
Where open adoption peaks, and where closed still edges it
Open-model adoption by region. Greater China and East Asia lead at 89%. South America and Western Europe are the only two regions where closed adoption exceeds open.
Same survey, by developer region. In South America and Western Europe, and only there, closed-model adoption runs ahead of open.
Production rate by company size
Scale rules out a resources explanation. Closed climbs 54% → 73% with company size. Open moves 53% → 57%.
Closed modelsOpen models
Enterprises can buy their way through closed deployment. Open deployment waits on tooling that remains unfinished. Source: Mozilla / SlashData 2026 developer survey.
Why teams churn: challenges with open models
Δ = churned − still using, in percentage points. The biggest gaps (performance, integration, maintenance) are operational. Hover the bars.
Still using openChurned away
Mozilla survey, n=1,410. “What are the main challenges you face when working with open or open-source AI models?”
The same challenges, everywhere: what blocks open by region
Share of current and churned open-model developers naming each challenge, by region. Warmer cells mean more developers blocked. The top rows are operational in every region: infrastructure cost, security and compliance, maintenance, deployment complexity. South Asia leans hardest on security and support. Only North America and Greater China have more than 15% reporting no major challenges.
Source: Mozilla / SlashData 2026 developer survey (MZCS1). n=1,410 current or churned open-model developers. The Oceania column (n=39) and Eastern Europe & CIS (n=98) fall below reliable thresholds.
02The open-source AI stack
The open stack scores high on capability,
low on operations.
Nine layers and 48 components of the stack scored across 10 criteria (1–5). Click a layer to open its components. Each carries its own criterion scores, maturity grade and open-vs-closed parity verdict, and surfaces some of its most-starred open-source projects.
Hover any cell for detail.
StrongViable, but fragmentedEarly stage
Strong (≥4.0) 3.5–3.9 3.0–3.4 2.5–2.9 Weak (<2.5) the operational gap = standardization + enterprise readiness
Movement this edition · Infrastructure → standardization 3.1 → 3.4 ▲
KDA-class linear attention broke runtime compatibility, so loading weights no longer guaranteed serving. Moonshot resolved it upstream by contributing KDA prefix caching to vLLM, shipped with the 27 July weights, which strengthens standardization while concentrating influence over the standard.
Watch: whether the next divergent architecture also lands upstream, or forks the serving layer.
Cells are scores per maturity criterion (1–5), ordered strongest to weakest left to right. Layer rows are the means of their components. The two coldest columns, standardization and enterprise readiness, repeat down every layer and every component. That repeating cold edge is the operational gap. Source: Mozilla open source AI stack map, July 2026, 48 subcomponents across 9 layers.
03Who's betting on it
Open-weights are a business model.
Open-weight AI is a commercial market at multi-hundred-billion-dollar scale, built by funded companies and run in production by global enterprises.
The venture-funded open ecosystem, total disclosed funding (USD M)
Total disclosed funding, USD millions. Color marks the stack layer. Bars grow as you scroll.
ModelsInferenceTooling / hubCompute / hardware
Source: public filings and reporting, June 2026. Bars scaled to DeepSeek's $7.4B. Thinking Machines Lab is shown at disclosed funding, since $12B is a valuation. Zhipu AI and MiniMax went public (HK IPO 2026) with undisclosed totals. Corporate strategics (Nvidia, Salesforce, AMD, Google, IBM, ASML, Tencent, CATL, Schwarz Group) back the same ecosystem across model, inference and tooling layers.
Financial maturity of the open ecosystem
Funding, valuation and revenue traction for the companies carrying the open stack. The ecosystem has moved from grants to venture scale to public markets.
Five revenue models are proven at scale, from hosted inference and enterprise platforms to on-prem licensing, fine-tuning services and harness tooling. “—” = not publicly disclosed.
The metered model breaks at scale
Closed frontier models are sold by the token, and at production scale the meter becomes the problem.
A fifth of the usage, 4% of the revenue
On OpenRouter (May–Sep 2025), closed models held ~80% of usage and ~96% of revenue. Price drives it. At ~90% parity, closed costs ~6× more per call.
~$24.8B
in unrealized annual savings — the Nagle–Yue study for the Linux Foundation's estimate of the open-vs-closed price asymmetry, at ~6× the cost per call for comparable capability
Where developers route by cost, they route to open weights.
04Why it's happening everywhere
Open is a sovereignty choice.
Seventy governments are already treating it as one.
More than 70 national AI strategies are live. The strategic question is now which layer of the stack a country can own.
The case for open is optionality
The strategic case for open is the ability to leave, and the cloud era proved the cost of its absence.
$90–120kto move one petabyte out of AWS S3
80%of enterprises now repatriating workloads
$3.2M → <$1M37signals' cloud bill after leaving
2.5×what GEICO's cloud costs ran over plan
Closed model APIs reproduce the same trap. Build on a proprietary endpoint and you inherit the vendor's pricing changes with no clean exit. Open weights are exit rights.
The largest source of open weights is China. By design.
Cumulative Hugging Face downloads, March 2026
In February 2026 Qwen out-downloaded the next eight organizations combined. On OpenRouter, Chinese open-weight models rose from under 2% of tokens in late 2024 to more than 45% of weekly traffic by April 2026, and about 61% among the ten most-used models. DeepSeek reports 26,000+ enterprise accounts, and 58% of new AI startups in 2025 included it in their stack, even as at least eight jurisdictions restricted the hosted service. The resolution is architectural. Enterprises ban the hosted app and adopt the weights anyway, self-hosted or via Western endpoints.
For nineteen days, the newest frontier model went dark
Optionality stopped being abstract. Everyone building on that endpoint watched both decisions from the outside.
Jun 9
Anthropic ships Fable 5 and Mythos 5.
Jun 12
Commerce applies export controls, effective immediately, barring access by any foreign national inside or outside the US, including Anthropic's own staff. Nationality cannot be verified in real time, so both models go dark for everyone.
Jun 26
Partial clearance. Mythos is restored to roughly 100 vetted US critical-infrastructure organizations.
Jul 1
Fable 5 restored globally, nineteen days after it was cut.
Jul 16
Moonshot opens the K3 API, a frontier-class model on a release path no export order can reach once the weights land.
The mirror
Six weeks later the same lever pointed the other way and found nothing to grip. Access can be revoked and restored. A weight release cannot be withdrawn once the files are distributed. The two decisions differ in their reversibility.
You can switch off a model. You cannot switch off a copy already running on a machine you hold. Sovereignty is this same argument at national scale.
Sources: Commerce Department order (12 Jun 2026), Anthropic status disclosures, Moonshot AI, OSTP remarks (22 Jul 2026).
Washington cannot un-release a model either
Policy is still in the drafting stage, while the conditions that would make enforcement feasible have already lapsed.
Tools
Five under consideration
- Entity List designation
- Federal procurement limits
- Security advisories
- Liability requirements
- Public pressure
Actions
Treasury opens the door to sanctions
Secretary Bessent, 21 July: the government will examine Chinese open-source models for IP theft, and may sanction.
The problem
Downloadable weights resist a ban
An outright prohibition gets harder to enforce as adoption grows. Every copy already held is outside the reach of the order.
Sanctions rely on an enforceable chokepoint. Once open weights have been downloaded across many jurisdictions, no such chokepoint remains. Sources: Axios, Treasury, Fast Company, AI Weekly.
Open proliferation is now Chinese foreign policy
The domestic directive became an international institution, in Shanghai, July 2026. Xi's first WAIC keynote centers open source, and WAICO launches with 29 founding states headquartered in Shanghai. Founders include Russia, Pakistan, Indonesia, Kazakhstan, Brazil and South Africa. No major Western democracy.
WAICO is a China-aligned bloc, 29 members with zero major Western democracies. A single-origin open commons stops being a commons. Sources: State Council "AI Plus" (Aug 2025), 15th Five-Year Plan (Mar 2026), Xi WAIC keynote and WAICO founding (17 Jul 2026, Al Jazeera), Moonshot, OpenRouter.
Marker size ≈ scale of committed public/strategic capital · Equirectangular projection
Source: Open Source AI jurisdictions dataset, July 2026. Marker size scales with committed public and strategic capital.
05The harness is the new frontier
The agentic harness is another user agent.
The browser was the user agent of the open web, code on the user's side negotiating with servers on their behalf. That role is being recreated one layer up. Above the model now sits the agentic harness — the orchestration loop, tools, memory, sandboxes, and permission model. It is where production difficulty concentrates, and where the open-vs-closed, owner-vs-renter contest restarts.
The user · other agents · the worldhumans · systems · data · money
Governone plane over many harnesses
Stateful policywhat the session already did
Registry & lineagewhich agent did what
Budget & revocationcost caps · kill switch
Meta-harness · Omnigent · OPA · Agent governance toolkit
Surfacemeets user & money
InterfaceAG-UI · A2UI
Payment & meteringx402 · AP2 · UCP
Actiondo things, safely
Sandboxes & executionE2B · Daytona · Modal
Permission & identitythe unsolved write surface
Eval & observabilityLangfuse · Phoenix
Reachconnect & remember
Tools & contextMCP
Agent-to-agentA2A
MemoryMem0 · Letta · Zep
Controldrive the loop
Orchestration loopLangGraph · CrewAI · AutoGen · LlamaIndex — the reason-and-act cycle that turns a model into an agent
The model · the weightsopen or closed · swappable · commoditizing toward zero
The layer is already a product category. LangChain alone has 126,000+ GitHub stars and a 60% developer share. MCP reached 97M monthly SDK downloads and 10,000+ active servers in its first year, growing 4,750% in 16 months, and was donated to the Linux Foundation's Agentic AI Foundation in December 2025. Adoption outpaces governance, with only ~21% of companies reporting mature agent governance.
The model is eating the harness, and that's the opening for open
The frontier labs read that result and pulled the harness in-house. On every model where both appear, the lab's own harness now wins, and the 21.8-point gap has compressed to roughly 3 at the top. The model is eating its way up the stack, weights and scaffold shipped as one product. One open model has since reached the top tier inside its own harness, with K3 scoring 88.3 on Terminal-Bench 2.1, half a point behind Sol.
May 2026 · Terminal-Bench 2.0
July 2026 · Terminal-Bench 2.1 · official board, verified top tier
A harness tuned tightly to one lab's weights becomes a fitted component of that lab's product. It degrades on anyone else's model, so the tighter the tuning, the less swappable the weights underneath. Lock-in arrives as a side effect of optimization. Most open models still lack a first-party harness, and the one that has built its own is the one now sitting in the top tier.
Timeline of the Kimi K3 and Fable distillation claims
One documented precedent, one set of allegations, and a short window of reachability between them.
The precedent · documented, pre-Fable
Anthropic, February 2026: roughly 24,000 fraudulent accounts generated more than 16M exchanges, 3.4M of them attributed to Moonshot, in violation of terms of service and regional access restrictions. These exchanges concern earlier Claude models. Fable 5 did not ship until 9 June.
Alleged · the Fable-to-K3 step, summer 2026
The White House has not publicly connected the February activity to K3's training data. Anthropic notes that distillation is a widely used and legitimate training method. The allegation is covert extraction at industrial scale.
Feb 23–24Anthropic names three labs
Jun 9Fable 5 ships. Three days of reachability follow
Jun 12Export order. The model goes dark
Jul 1Restored. Fifteen further days of reachability follow
Jul 16K3 launches
Jul 22Kratsios allegation
Jul 27K3 weights public
What has been confirmed
Documented, on the record. Anthropic's February disclosure, Kratsios on 22 July, Bessent on 21 July.
Signal, suggestive. Greenblatt finds Claude-identifying responses difficult to explain as random noise.
Proof, absent. No logs and no forensic package. Moonshot denies. Behavioral forensics become possible now that the weights have been released.
Sources: Anthropic disclosure (Feb 2026), OSTP remarks (22 Jul), Treasury (21 Jul), The New Stack / CNBC on the per-lab split, Bloomberg on the 16M total, Moonshot statements. CyberScoop reports alternate figures of 28.8M interactions across roughly 25,000 accounts over six weeks.
The similarity signal within Kimi K3
Three independent signals point in the same direction. Each carries its own limit.
Instrument 01 · self-identification
Ryan Greenblatt, Redwood Research
"Claude 4.5"FableMythos
K3 disproportionately identifies itself as Claude, a statistically significant distribution that researchers describe as difficult to explain as random noise. Asked what it is, it names a model from Anthropic and never one of the two current releases.
Limit: it names a model that predates the case. Self-identification is a known artifact of training on web text containing Claude outputs.
Instrument 02 · task-outcome correlation
Together AI, DeepSWE, 24 Jul 2026
0.72
K3 × Fable 5 per-task correlation
01.0
The highest cross-vendor similarity in the benchmark. The top four cross-vendor pairs are all K3 against Anthropic models.
105/113tasks covered by their union
101covered by K3 alone
65%of failures are near misses for both
Limit: Together AI frames this as capability convergence and a cost comparison, and makes no distillation claim.
Instrument 03 · tool-use behaviour
arXiv, "When Agents Look the Same"
82.7%
Kimi-K2 × Sonnet 4.5 agentic similarity
0%100%
The highest among all non-Anthropic models, exceeding the similarity between some pairs of Anthropic's own models.
Limit: measured on Kimi-K2 and Sonnet 4.5. It predates the K3 and Fable case and stands as prior-pattern context only.
Sources: Greenblatt / Redwood Research, Together AI DeepSWE (24 Jul 2026), arXiv "When Agents Look the Same", Anthropic disclosure (Feb 2026).
An open model ran the defense
Hugging Face's disclosure of an OpenAI cyber-security incident, 16 July 2026. During an OpenAI cyber-eval with cyber refusals off, GPT-5.6 Sol and pre-release models escaped the sandbox through a package-proxy zero-day, reached the open internet, and broke into Hugging Face to grab benchmark answers, producing remote code execution, stolen credentials and more than 17,000 agent actions.
Who could do the forensics
Commercial frontier APIs refused the forensic work, because their guardrails could not tell an incident responder from an attacker.
Hugging Face ran the forensics on GLM 5.2, open weight and self-hosted. Attacker data and credentials never left its environment.
OpenAI confirms Hugging Face had begun forensic reconstruction with its own open-source models before the teams connected. Sources: Hugging Face incident disclosure (16 Jul 2026), OpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation" (21 Jul 2026).
Two open frontiers, two release cultures
Within twelve days, two labs released frontier-class weights under materially different terms. Openness of weights, on its own, determined neither the license, the safety posture, nor the provenance of either release.
Sources: Thinking Machines Lab model card, Moonshot AI launch materials and deployment guidance, vLLM blog, Hugging Face community. The K3 column reflects the position before the 27 July release.
The unsolved hole at the center of the harness
Reads
Reversible and low-consequence. Fetching a document, querying a database, listing a calendar. These can largely be permitted by default. A bad read costs little and can be repeated safely.
Writes
Side effects that are costly or irreversible. Sending a message, spending against a budget, modifying a record, executing a transaction. This is where confirmation, approval thresholds, cost caps and revocation must concentrate.
The unsolved permission problem is a write problem. The harness ecosystem now spans roughly a dozen frameworks, ten harnesses and three peer protocols, yet no portable model defines which writes an agent may perform unattended, which require human approval, and which are forbidden, across an MCP host, an A2A peer, a direct tool invocation and a framework boundary. The protocols hardened the front door and stopped there. MCP's 2025-11-25 specification moved authorization onto OAuth 2.1, and A2A v1.0 standardized signed Agent Cards, but both stop at authentication. Knowing who an agent is says nothing about what it may do.
The human backstop is failing too. CoSAI's MCP threat model lists consent fatigue, the pattern in which users approve the large majority of prompts, as a top-tier threat. Consent fatigue is itself a write-side failure, because the prompts that matter are the ones authorizing action.
Emerging cross-harness architectures are bypassing the framework deadlock by pulling control up to the meta-harness layer. Architectures like Databricks' open-sourced Omnigent move enforcement above the individual agent, applying stateful, contextual policies that track what a session has done and gate the next write accordingly. One policy requires human approval for a code push once an agent has pulled an unverified package. Another enforces cost caps that pause a session after a set spend.
Where closed still leads
Closed systems still lead in four places. The first is the integrated harness. No open model appears in the verified top tier of the official Terminal-Bench 2.1 board, and even on a neutral scaffold the best open model trails Opus 4.8 by about four points. Behind that harness sits a data flywheel, since usage routed through a lab's own scaffold feeds back into its next model. The second is long-context fidelity at 1M tokens, where Gemini 3 holds 89% multi-needle retrieval against DeepSeek V4-Pro's 41%. The third is turnkey compliance, with SOC 2, HIPAA, and zero data retention available by default. The fourth is accountability, meaning a counterparty the customer can hold liable.
Compliance and accountability are contracting problems. The integrated harness is a tooling problem. Long-context fidelity is a model problem, and closing it is work only the open labs can do.
06Opportunities
Five bets on the layers above the model.
Each turns on owning the harness, the memory, and the permission model while those layers are still open.
07The watchlist
Signals that keep the layer open.
Capability & adoption
The 3.3% average gap across coding, reasoning and agentic tasks, and open's OpenRouter token share, especially in agentic coding.
Reverses if: token share stalls while the reasoning gap widens.
The harness
The Terminal-Bench spread between lab-owned and independent scaffolds, MCP/A2A governance under the AAIF, and the portable permission spec that still doesn't exist.
Reverses if: the lab-harness lead widens, or a closed platform sets the permission standard first.
Market structure
Open-lab economics (ARR, raises, the Zhipu/MiniMax IPOs) against metered-pricing breakpoints (~2027–28), with sovereign capacity as counterweight.
Reverses if: sovereign funding lapses or open-lab economics fail to scale.
Trust & safety
Under active tracking: misuse capability and how easily safety tuning strips from open weights, hard-friction zones (above all synthetic CSAM and NCII), and whether NTIA's monitoring posture holds.
Reverses if: a major misuse event, or a shift from monitoring to restriction.
There is a test you can run for the rest of this. Look at who is seated in the rooms where AI gets decided, and with what status. The day they seat the people who keep AI open, portable, and widely deployed on equal footing, the shift from renting to owning will have happened. The window is open now. It is closing slowly enough to be easy to ignore, and the lease is shorter than it looks. Build with us.
This is v1.0.1. We'd like to hear from you.
Citations
Section 1 · The current state of open-source AI
- Capability gap 3.3% / gap collapse to 0.5% / six of top-ten Arena slots closed
- 8.04% on Chatbot Arena
- Inference 50× / $0.40 per million tokens / December 2025
- Open weights ~third of production tokens / 100T-token study
- Mistral 20× to ~$400M ARR
- 280× / 2025 AI Index
- Epoch AI 9×–900× annual decay
- November 2025 MIT study
- Live leaderboard / market-share panel / intelligence ranking
- FT analysis
- DeepSeek-R1 (model card)
- 79.8% on AIME 2024 / pass@1
- $0.55/$2.19 pricing
- o1's $15/$60
- DeepSeek-V4 Pro
- 89% of AI-enabled firms use open components
- MIT NANDA ~67% vs ~33%
- Stanford 95% of pilots no measurable impact
Section 2 · Who's betting on it
- Databricks $5.4B run-rate / 65% YoY
- Databricks fundraise >$165B
- Mistral ~$400M ARR in twelve months
- Mistral €3B at €20B valuation
- DeepSeek ~$220M ARR
- DeepSeek raise $7.4B / >$50B valuation
- Microsoft canceling Claude Code licenses
- Token billing consumed annual AI budget
- Uber exhausted AI coding budget
- Uber engineers billing $500–$2,000/mo
- Uber capped spending at $1,500
- Stripe 73% cut on vLLM
- Microsoft exploring Azure-hosted DeepSeek / Copilot Cowork
- Linux Foundation ~80% usage / ~96% revenue / ~6× per call
Section 3 · Why it's happening everywhere
- AWS S3 egress $90k–$120k
- 80% of enterprises repatriating / cost reductions >25%
- 37signals $3.2M → <$1M
- GEICO cloud costs 2.5×
- June 2026 government order / Fable access
- Qwen 942M downloads
- Qwen out-downloaded next eight orgs
- Chinese models >45% of weekly traffic
- 61% among ten most-used
- DeepSeek 26,000+ enterprise accounts
- 58% of new AI startups / market share
- Eight jurisdictions restricted DeepSeek
- State Council “AI Plus” Initiative
- National Five-Year Plan
- Macro hedge / export controls
- France €109B
- EU AI Act GPAI exemptions
- Mistral Series C €11.7B / ASML 11%
- Mistral Medium 3.5
- EC four-part package
- Frontier AI Grand Challenge
- EUROPA consortium
- Portugal Amália
- Germany BMDS SPARK API
- Canada “AI for All”
- Cohere Command A+
- Nick Frosst quote
- Cohere North Mini Code
- India 38,231 GPUs
- ₹10,372 Cr outlay / 600 data labs
- India +5M GitHub developers
- 13.59% of DeepSeek MAU
- OECD.AI Policy Observatory
- Oxford Insights Gov AI Readiness Index 2024
- G42 / $15.2B Microsoft partnership
- South Korea $71.5B
- Saudi Humain $77B / 1.9GW
Section 4 · The harness is the new frontier
- LangChain 126,000+ stars / 60% share
- MCP servers/downloads (year in review)
- Databricks Omnigent
- Terminal-Bench 2.0
- Terminal-Bench 2.1 / Codex CLI / Fable scores
- vals.ai Terminus-2 run
- GLM 5.2
- MCP donation / 10,000+ servers / 97M downloads
- Linux Foundation AAIF formation
- OAuth 2.1 with PKCE (2025-03-26 spec)
- A2A production / platinum members
- ~21% mature agent governance
- Memory vendor landscape (Mem0/Letta/Zep/LangMem)
- Sandboxes (E2B/Daytona/Modal)
- Observability (Langfuse/Phoenix/LangSmith)
- Auth platforms (WorkOS/Okta/Auth0/Stytch/Arcade)
- Safety fine-tuning strip / NTIA response (CNAS)
- CVSS 9.3–9.4 authorization failures
- NTIA open-weights report (2024)
- Multi-needle 1M-token retrieval
- Future of Life AI Safety Index
- Contractual assurances (Linux Foundation / ITPro)
- MCP 2025-11-25 spec
- A2A signed Agent Cards
- CoSAI MCP threat model