The State of Open Source AI — v1.0.1 · July 2026

26 min read Original article ↗

An introduction from our CTO

The Māori language Te Reo has no commercial market to speak of, but a broadcaster in New Zealand's far north has been building speech models for it anyway — and all under a license that keeps the recordings with the communities that originally gave them. Half a world away, in a cellular dead zone in East Africa, farmers point a phone at a cassava leaf and get a diagnosis from a model small enough to live on the handset. Neither project needed permission. And neither could have been rented from a frontier model.

The largest companies on earth are now thinking this way. This month, on the Artificial Analysis Intelligence Index, the strongest closed model scored 61. The strongest open model? 57. The open model came in fourth overall, ahead of models from three of the biggest closed labs. At the end of 2025, about a third of all tokens on OpenRouter were being routed to open-weight models; now the seven highest-volume models on the platform all ship open weights.

In a surprising recent turn, corporate heavyweights are leaning in an encouragingly unexpected direction: On July 24, Mozilla signed an open letter alongside NVIDIA, Microsoft, Meta, IBM, Dell, Mistral, Hugging Face, the Linux Foundation, and Andreessen Horowitz that argued that open weights are too important to be ignored.

We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community made sure it never could. We bet on open the first time. Open won. Together, we can do it again.

Raffi Krikorian · Chief Technology Officer, Mozilla

01The current state of open models

The model layer has commoditized.
Value accrues to the harness above it.

Open weights are where the work happens. A majority of production tokens now route through them, and the seven highest-volume models on OpenRouter are all open weight. Closed models still lead at the frontier, on reasoning and multimodality. However, most production workloads run well below that ceiling. Commodity inputs surrender pricing power. Value moves up to the agentic harness, the layer above the model.

Terminology note

Open model here means weights you can download, run and modify on hardware you control. That category splits. Open weights means the parameters ship under a permissive license with no training code and no data documentation, which describes most of what this report measures. Open source AI, as OSI defines it, additionally requires the training code and enough information about the data to rebuild the system.

The frontier's top three are closed. The fourth is open.

Artificial Analysis Intelligence Index v4.1, a nine-eval composite that includes Terminal-Bench 2.1, Humanity's Last Exam and GPQA Diamond. Kimi K3 ranks fourth, four points off the closed frontier.

Open weightsClosed

4 ptsfrom the best open model (Kimi K3, 57) to the closed frontier (Claude Opus 5, 61)

3 of 11top-tier models shipping open weights

Source: Artificial Analysis Intelligence Index v4.1, July 2026, top 11 of 586 models shown. Open = downloadable weights.

3.6 points off the top, at a third of the price

The same index against weighted-average price per 1M input tokens, for the top ten. Kimi K3 is the only open-weight model in the leading cluster, fourth overall and 3.6 points off the top, at about a third of the price.

Source: Artificial Analysis Intelligence Index v4.1 via OpenRouter, July 2026. Price per 1M input plotted where list price is disclosed.

Open trails the frontier by 6 ECI points, about one release cycle

Epoch Capabilities Index by release date. The best open model, Kimi K3 at 156, against the closed frontier, GPT-5.6 Sol at 162.

Source: Epoch AI, Epoch Capabilities Index (CC-BY), July 2026. Open = downloadable weights. The gap is the point difference at the frontier, and the confidence intervals overlap.

A jagged frontier: parity, contested, a closed edge

For most workloads open models already clear the bar. The closed frontier earns its premium in a narrow band, namely deep reasoning, long-context reliability and professional-knowledge polish. Match the model to the job and you need the frontier for less than you think.

Open leads or parity

Frontend coding

K3 debuted first on LMArena's Frontend Code Arena at 1679 Elo, leading in six of seven frontend domains, plus coding, instruction-following and general knowledge.

Frontend Code Arena measures building website and UI code, scored by blind developer votes.

Contested

Agentic terminal work

K3 scores 88.3 against Sol's 88.8 on Terminal-Bench 2.1. It wins Program Bench, SpreadsheetBench 2 and BrowseComp, and loses FrontierSWE at 81.2 against Fable 5's 86.6.

Terminal-Bench, Program Bench and BrowseComp measure agent work, running terminal tasks, writing programs and researching the web. FrontierSWE measures resolving real software-engineering tickets.

Closed edge

Professional knowledge work

Fable 5 leads K3 by 92 Elo on GDPval-AA v2, the largest Elo separation among the shared benchmarks, alongside long-context fidelity and conversational polish, which Moonshot concedes still trails.

GDPval-AA v2 measures expert-graded professional knowledge work, where long-context reliability and polish appear.

Scores use different scales and come mostly from vendor-run tests, so treat them as directional. LMArena uses blind human voting. Sources: LMArena, Artificial Analysis, Moonshot K3 blog.

Inference fell 50× in 36 months

Cheapest model at GPT-4-class performance, blended API list price, log scale, against prior platform-shift cost curves at their historical rates. Over the same 36 months the dotcom bandwidth curve delivers 2.6× and the PC compute curve 3.4×. The frontier price fell 112× from GPT-4's $45 launch.

Sources: a16z "LLMflation" methodology, Epoch AI LLM inference price trends, Artificial Analysis, Moonshot, MTS/Substack.

Open weights win the tokens

The share of tokens routed on OpenRouter through open-weight models grew from a negligible base to a third by late 2025 to a majority by mid-2026.

Source: OpenRouter 100T-token study (Nov 2024–Nov 2025) and live leaderboard, with intermediate points interpolated. By request count, closed US providers still lead. The open lead is a token-volume lead, concentrated in coding and agentic workloads.

Open weights dominate token usage

Each model's share of the top-20 routed token volume on OpenRouter, 1–27 July 2026. The top seven by volume are all open weight. Anthropic's closed Claude models are the next entrants.

Open weightsClosed

Top 20 tokens, by licence

Open, ranks 1–10 · 72.4% 8.7% 18.9%

72.4% Open, ranks 1–10

8.7% Closed, ranks 1–10

18.9% Ranks 11–20

K3 note: the API opened 16 July and the weights on 27 July, so K3 does not appear in July's routed-token top 20. New subscriptions were temporarily paused on 20 July as demand neared capacity. These ten account for 81.1% of top-20 volume, leaving 18.9% across the remaining ten. Shares are of the top 20 only, not of all OpenRouter traffic. Source: OpenRouter LLM Leaderboard, July 2026 (this month), routed traffic only.

The open frontier and the deployable open frontier split

Minimum serving footprint per model, measured in 8-GPU nodes. Bar length shows nodes to run, and memory is the figure that produces it. Frontier capability is downloadable. Node count decides who can run it.

Node counts follow Moonshot's guidance of 64+ accelerators for K3. A single 8-GPU node barely holds K3's weights, so eight nodes reflects serving it rather than only loading it. Inkling's ≥600 GB straddles the one-node line depending on GPU generation. K3 open weights released 27 July 2026. Sources: Moonshot deployment guidance, TML model card, Hugging Face community, Northflank.

Open ships easy.
Open deploys hard.

Data from the Mozilla / SlashData 2026 developer survey. Open models lead in adoption. 79% of developers adding AI functionality use them, against 71% for closed, and the two are largely complementary, with half of developers using both. But production is where teams stall. Only 53% of open-model teams reach production versus 63% for closed. The gap traces to operational tooling and trust.

Open models lead in adoption, and mostly coexist with closed

Share of developers adding AI functionality to their applications who currently use each model type, and how the two overlap.

How they combine

29%OS only

50%Both

21%CS only

Source: Mozilla / SlashData 2026 developer survey. Most teams treat open and closed as complements, with 50% running both, 29% open only and 21% closed only.

Where open adoption peaks, and where closed still edges it

Open-model adoption by region. Greater China and East Asia lead at 89%. South America and Western Europe are the only two regions where closed adoption exceeds open.

Same survey, by developer region. In South America and Western Europe, and only there, closed-model adoption runs ahead of open.

Production rate by company size

Scale rules out a resources explanation. Closed climbs 54% → 73% with company size. Open moves 53% → 57%.

Closed modelsOpen models

Enterprises can buy their way through closed deployment. Open deployment waits on tooling that remains unfinished. Source: Mozilla / SlashData 2026 developer survey.

Why teams churn: challenges with open models

Δ = churned − still using, in percentage points. The biggest gaps (performance, integration, maintenance) are operational. Hover the bars.

Still using openChurned away

Mozilla survey, n=1,410. “What are the main challenges you face when working with open or open-source AI models?”

The same challenges, everywhere: what blocks open by region

Share of current and churned open-model developers naming each challenge, by region. Warmer cells mean more developers blocked. The top rows are operational in every region: infrastructure cost, security and compliance, maintenance, deployment complexity. South Asia leans hardest on security and support. Only North America and Greater China have more than 15% reporting no major challenges.

Source: Mozilla / SlashData 2026 developer survey (MZCS1). n=1,410 current or churned open-model developers. The Oceania column (n=39) and Eastern Europe & CIS (n=98) fall below reliable thresholds.


02The open-source AI stack

The open stack scores high on capability,
low on operations.

Nine layers and 48 components of the stack scored across 10 criteria (1–5). Click a layer to open its components. Each carries its own criterion scores, maturity grade and open-vs-closed parity verdict, and surfaces some of its most-starred open-source projects.

Hover any cell for detail.

StrongViable, but fragmentedEarly stage

Strong (≥4.0) 3.5–3.9 3.0–3.4 2.5–2.9 Weak (<2.5) the operational gap = standardization + enterprise readiness

Movement this edition · Infrastructure → standardization 3.1 → 3.4 ▲

KDA-class linear attention broke runtime compatibility, so loading weights no longer guaranteed serving. Moonshot resolved it upstream by contributing KDA prefix caching to vLLM, shipped with the 27 July weights, which strengthens standardization while concentrating influence over the standard.

Watch: whether the next divergent architecture also lands upstream, or forks the serving layer.

Cells are scores per maturity criterion (1–5), ordered strongest to weakest left to right. Layer rows are the means of their components. The two coldest columns, standardization and enterprise readiness, repeat down every layer and every component. That repeating cold edge is the operational gap. Source: Mozilla open source AI stack map, July 2026, 48 subcomponents across 9 layers.


03Who's betting on it

Open-weights are a business model.

Open-weight AI is a commercial market at multi-hundred-billion-dollar scale, built by funded companies and run in production by global enterprises.

The venture-funded open ecosystem, total disclosed funding (USD M)

Total disclosed funding, USD millions. Color marks the stack layer. Bars grow as you scroll.

ModelsInferenceTooling / hubCompute / hardware

Source: public filings and reporting, June 2026. Bars scaled to DeepSeek's $7.4B. Thinking Machines Lab is shown at disclosed funding, since $12B is a valuation. Zhipu AI and MiniMax went public (HK IPO 2026) with undisclosed totals. Corporate strategics (Nvidia, Salesforce, AMD, Google, IBM, ASML, Tencent, CATL, Schwarz Group) back the same ecosystem across model, inference and tooling layers.

Financial maturity of the open ecosystem

Funding, valuation and revenue traction for the companies carrying the open stack. The ecosystem has moved from grants to venture scale to public markets.

Five revenue models are proven at scale, from hosted inference and enterprise platforms to on-prem licensing, fine-tuning services and harness tooling. “—” = not publicly disclosed.

The metered model breaks at scale

Closed frontier models are sold by the token, and at production scale the meter becomes the problem.

A fifth of the usage, 4% of the revenue

On OpenRouter (May–Sep 2025), closed models held ~80% of usage and ~96% of revenue. Price drives it. At ~90% parity, closed costs ~6× more per call.

~$24.8B

in unrealized annual savings — the Nagle–Yue study for the Linux Foundation's estimate of the open-vs-closed price asymmetry, at ~6× the cost per call for comparable capability

Where developers route by cost, they route to open weights.


04Why it's happening everywhere

Open is a sovereignty choice.
Seventy governments are already treating it as one.

More than 70 national AI strategies are live. The strategic question is now which layer of the stack a country can own.

The case for open is optionality

The strategic case for open is the ability to leave, and the cloud era proved the cost of its absence.

$90–120kto move one petabyte out of AWS S3

80%of enterprises now repatriating workloads

$3.2M → <$1M37signals' cloud bill after leaving

2.5×what GEICO's cloud costs ran over plan

Closed model APIs reproduce the same trap. Build on a proprietary endpoint and you inherit the vendor's pricing changes with no clean exit. Open weights are exit rights.

The largest source of open weights is China. By design.

Cumulative Hugging Face downloads, March 2026

In February 2026 Qwen out-downloaded the next eight organizations combined. On OpenRouter, Chinese open-weight models rose from under 2% of tokens in late 2024 to more than 45% of weekly traffic by April 2026, and about 61% among the ten most-used models. DeepSeek reports 26,000+ enterprise accounts, and 58% of new AI startups in 2025 included it in their stack, even as at least eight jurisdictions restricted the hosted service. The resolution is architectural. Enterprises ban the hosted app and adopt the weights anyway, self-hosted or via Western endpoints.

For nineteen days, the newest frontier model went dark

Optionality stopped being abstract. Everyone building on that endpoint watched both decisions from the outside.

Jun 9

Anthropic ships Fable 5 and Mythos 5.

Jun 12

Commerce applies export controls, effective immediately, barring access by any foreign national inside or outside the US, including Anthropic's own staff. Nationality cannot be verified in real time, so both models go dark for everyone.

Jun 26

Partial clearance. Mythos is restored to roughly 100 vetted US critical-infrastructure organizations.

Jul 1

Fable 5 restored globally, nineteen days after it was cut.

Jul 16

Moonshot opens the K3 API, a frontier-class model on a release path no export order can reach once the weights land.

The mirror

Six weeks later the same lever pointed the other way and found nothing to grip. Access can be revoked and restored. A weight release cannot be withdrawn once the files are distributed. The two decisions differ in their reversibility.

You can switch off a model. You cannot switch off a copy already running on a machine you hold. Sovereignty is this same argument at national scale.

Sources: Commerce Department order (12 Jun 2026), Anthropic status disclosures, Moonshot AI, OSTP remarks (22 Jul 2026).

Washington cannot un-release a model either

Policy is still in the drafting stage, while the conditions that would make enforcement feasible have already lapsed.

Tools

Five under consideration

  • Entity List designation
  • Federal procurement limits
  • Security advisories
  • Liability requirements
  • Public pressure

Actions

Treasury opens the door to sanctions

Secretary Bessent, 21 July: the government will examine Chinese open-source models for IP theft, and may sanction.

The problem

Downloadable weights resist a ban

An outright prohibition gets harder to enforce as adoption grows. Every copy already held is outside the reach of the order.

Sanctions rely on an enforceable chokepoint. Once open weights have been downloaded across many jurisdictions, no such chokepoint remains. Sources: Axios, Treasury, Fast Company, AI Weekly.

Open proliferation is now Chinese foreign policy

The domestic directive became an international institution, in Shanghai, July 2026. Xi's first WAIC keynote centers open source, and WAICO launches with 29 founding states headquartered in Shanghai. Founders include Russia, Pakistan, Indonesia, Kazakhstan, Brazil and South Africa. No major Western democracy.

INSTITUTIONAL ARM · WAICO · SHANGHAI, JULY 2026 29 founding states · no major Western democracy SUPPLY Codified directive AI-Plus directive and the 15th Five-Year Plan make open-source proliferation a core state objective. Macro hedge against chip export controls. MECHANISM Release the weights Inference moves onto end users' own hardware, worldwide. No serving cost. No export surface. No chokepoint to sanction. DEMAND Adoption at both ends The Global South diversifies away from US tech. Well-capitalised firms, Microsoft among them, adopt for cost per task. 46.4% of routed tokens, against 35.7% US. JULY EXHIBIT Kimi K3 Billed as the world's first open 3T-class model, announced at WAIC before Xi's speech. Weights published 27 July. DISTRIBUTION RAILS 5,000 training slots over five years Pledged to developing countries, with cooperation centres planned with ASEAN, the Arab League, the African Union and BRICS. A single-origin open commons stops being a commons.

WAICO is a China-aligned bloc, 29 members with zero major Western democracies. A single-origin open commons stops being a commons. Sources: State Council "AI Plus" (Aug 2025), 15th Five-Year Plan (Mar 2026), Xi WAIC keynote and WAICO founding (17 Jul 2026, Al Jazeera), Moonshot, OpenRouter.

Marker size ≈ scale of committed public/strategic capital · Equirectangular projection

Source: Open Source AI jurisdictions dataset, July 2026. Marker size scales with committed public and strategic capital.


05The harness is the new frontier

The agentic harness is another user agent.

The browser was the user agent of the open web, code on the user's side negotiating with servers on their behalf. That role is being recreated one layer up. Above the model now sits the agentic harness — the orchestration loop, tools, memory, sandboxes, and permission model. It is where production difficulty concentrates, and where the open-vs-closed, owner-vs-renter contest restarts.

The user · other agents · the worldhumans · systems · data · money

Governone plane over many harnesses

Stateful policywhat the session already did

Registry & lineagewhich agent did what

Budget & revocationcost caps · kill switch

Meta-harness · Omnigent · OPA · Agent governance toolkit

Surfacemeets user & money

InterfaceAG-UI · A2UI

Payment & meteringx402 · AP2 · UCP

Actiondo things, safely

Sandboxes & executionE2B · Daytona · Modal

Permission & identitythe unsolved write surface

Eval & observabilityLangfuse · Phoenix

Reachconnect & remember

Tools & contextMCP

Agent-to-agentA2A

MemoryMem0 · Letta · Zep

Controldrive the loop

Orchestration loopLangGraph · CrewAI · AutoGen · LlamaIndex — the reason-and-act cycle that turns a model into an agent

The model · the weightsopen or closed · swappable · commoditizing toward zero

The layer is already a product category. LangChain alone has 126,000+ GitHub stars and a 60% developer share. MCP reached 97M monthly SDK downloads and 10,000+ active servers in its first year, growing 4,750% in 16 months, and was donated to the Linux Foundation's Agentic AI Foundation in December 2025. Adoption outpaces governance, with only ~21% of companies reporting mature agent governance.

The model is eating the harness, and that's the opening for open

The frontier labs read that result and pulled the harness in-house. On every model where both appear, the lab's own harness now wins, and the 21.8-point gap has compressed to roughly 3 at the top. The model is eating its way up the stack, weights and scaffold shipped as one product. One open model has since reached the top tier inside its own harness, with K3 scoring 88.3 on Terminal-Bench 2.1, half a point behind Sol.

May 2026 · Terminal-Bench 2.0

July 2026 · Terminal-Bench 2.1 · official board, verified top tier

A harness tuned tightly to one lab's weights becomes a fitted component of that lab's product. It degrades on anyone else's model, so the tighter the tuning, the less swappable the weights underneath. Lock-in arrives as a side effect of optimization. Most open models still lack a first-party harness, and the one that has built its own is the one now sitting in the top tier.

Timeline of the Kimi K3 and Fable distillation claims

One documented precedent, one set of allegations, and a short window of reachability between them.

The precedent · documented, pre-Fable

Anthropic, February 2026: roughly 24,000 fraudulent accounts generated more than 16M exchanges, 3.4M of them attributed to Moonshot, in violation of terms of service and regional access restrictions. These exchanges concern earlier Claude models. Fable 5 did not ship until 9 June.

Alleged · the Fable-to-K3 step, summer 2026

The White House has not publicly connected the February activity to K3's training data. Anthropic notes that distillation is a widely used and legitimate training method. The allegation is covert extraction at industrial scale.

Feb 23–24Anthropic names three labs

Jun 9Fable 5 ships. Three days of reachability follow

Jun 12Export order. The model goes dark

Jul 1Restored. Fifteen further days of reachability follow

Jul 16K3 launches

Jul 22Kratsios allegation

Jul 27K3 weights public

What has been confirmed

Documented, on the record. Anthropic's February disclosure, Kratsios on 22 July, Bessent on 21 July.

Signal, suggestive. Greenblatt finds Claude-identifying responses difficult to explain as random noise.

Proof, absent. No logs and no forensic package. Moonshot denies. Behavioral forensics become possible now that the weights have been released.

Sources: Anthropic disclosure (Feb 2026), OSTP remarks (22 Jul), Treasury (21 Jul), The New Stack / CNBC on the per-lab split, Bloomberg on the 16M total, Moonshot statements. CyberScoop reports alternate figures of 28.8M interactions across roughly 25,000 accounts over six weeks.

The similarity signal within Kimi K3

Three independent signals point in the same direction. Each carries its own limit.

Instrument 01 · self-identification

Ryan Greenblatt, Redwood Research

"Claude 4.5"FableMythos

K3 disproportionately identifies itself as Claude, a statistically significant distribution that researchers describe as difficult to explain as random noise. Asked what it is, it names a model from Anthropic and never one of the two current releases.

Limit: it names a model that predates the case. Self-identification is a known artifact of training on web text containing Claude outputs.

Instrument 02 · task-outcome correlation

Together AI, DeepSWE, 24 Jul 2026

0.72

K3 × Fable 5 per-task correlation

01.0

The highest cross-vendor similarity in the benchmark. The top four cross-vendor pairs are all K3 against Anthropic models.

105/113tasks covered by their union

101covered by K3 alone

65%of failures are near misses for both

Limit: Together AI frames this as capability convergence and a cost comparison, and makes no distillation claim.

Instrument 03 · tool-use behaviour

arXiv, "When Agents Look the Same"

82.7%

Kimi-K2 × Sonnet 4.5 agentic similarity

0%100%

The highest among all non-Anthropic models, exceeding the similarity between some pairs of Anthropic's own models.

Limit: measured on Kimi-K2 and Sonnet 4.5. It predates the K3 and Fable case and stands as prior-pattern context only.

Sources: Greenblatt / Redwood Research, Together AI DeepSWE (24 Jul 2026), arXiv "When Agents Look the Same", Anthropic disclosure (Feb 2026).

An open model ran the defense

Hugging Face's disclosure of an OpenAI cyber-security incident, 16 July 2026. During an OpenAI cyber-eval with cyber refusals off, GPT-5.6 Sol and pre-release models escaped the sandbox through a package-proxy zero-day, reached the open internet, and broke into Hugging Face to grab benchmark answers, producing remote code execution, stolen credentials and more than 17,000 agent actions.

Who could do the forensics

Commercial frontier APIs refused the forensic work, because their guardrails could not tell an incident responder from an attacker.

Hugging Face ran the forensics on GLM 5.2, open weight and self-hosted. Attacker data and credentials never left its environment.

OpenAI confirms Hugging Face had begun forensic reconstruction with its own open-source models before the teams connected. Sources: Hugging Face incident disclosure (16 Jul 2026), OpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation" (21 Jul 2026).

Two open frontiers, two release cultures

Within twelve days, two labs released frontier-class weights under materially different terms. Openness of weights, on its own, determined neither the license, the safety posture, nor the provenance of either release.

Sources: Thinking Machines Lab model card, Moonshot AI launch materials and deployment guidance, vLLM blog, Hugging Face community. The K3 column reflects the position before the 27 July release.

The unsolved hole at the center of the harness

Reads

Reversible and low-consequence. Fetching a document, querying a database, listing a calendar. These can largely be permitted by default. A bad read costs little and can be repeated safely.

Writes

Side effects that are costly or irreversible. Sending a message, spending against a budget, modifying a record, executing a transaction. This is where confirmation, approval thresholds, cost caps and revocation must concentrate.

The unsolved permission problem is a write problem. The harness ecosystem now spans roughly a dozen frameworks, ten harnesses and three peer protocols, yet no portable model defines which writes an agent may perform unattended, which require human approval, and which are forbidden, across an MCP host, an A2A peer, a direct tool invocation and a framework boundary. The protocols hardened the front door and stopped there. MCP's 2025-11-25 specification moved authorization onto OAuth 2.1, and A2A v1.0 standardized signed Agent Cards, but both stop at authentication. Knowing who an agent is says nothing about what it may do.

The human backstop is failing too. CoSAI's MCP threat model lists consent fatigue, the pattern in which users approve the large majority of prompts, as a top-tier threat. Consent fatigue is itself a write-side failure, because the prompts that matter are the ones authorizing action.

Emerging cross-harness architectures are bypassing the framework deadlock by pulling control up to the meta-harness layer. Architectures like Databricks' open-sourced Omnigent move enforcement above the individual agent, applying stateful, contextual policies that track what a session has done and gate the next write accordingly. One policy requires human approval for a code push once an agent has pulled an unverified package. Another enforces cost caps that pause a session after a set spend.

Where closed still leads

Closed systems still lead in four places. The first is the integrated harness. No open model appears in the verified top tier of the official Terminal-Bench 2.1 board, and even on a neutral scaffold the best open model trails Opus 4.8 by about four points. Behind that harness sits a data flywheel, since usage routed through a lab's own scaffold feeds back into its next model. The second is long-context fidelity at 1M tokens, where Gemini 3 holds 89% multi-needle retrieval against DeepSeek V4-Pro's 41%. The third is turnkey compliance, with SOC 2, HIPAA, and zero data retention available by default. The fourth is accountability, meaning a counterparty the customer can hold liable.

Compliance and accountability are contracting problems. The integrated harness is a tooling problem. Long-context fidelity is a model problem, and closing it is work only the open labs can do.


06Opportunities

Five bets on the layers above the model.

Each turns on owning the harness, the memory, and the permission model while those layers are still open.


07The watchlist

Signals that keep the layer open.

Capability & adoption

The 3.3% average gap across coding, reasoning and agentic tasks, and open's OpenRouter token share, especially in agentic coding.

Reverses if: token share stalls while the reasoning gap widens.

The harness

The Terminal-Bench spread between lab-owned and independent scaffolds, MCP/A2A governance under the AAIF, and the portable permission spec that still doesn't exist.

Reverses if: the lab-harness lead widens, or a closed platform sets the permission standard first.

Market structure

Open-lab economics (ARR, raises, the Zhipu/MiniMax IPOs) against metered-pricing breakpoints (~2027–28), with sovereign capacity as counterweight.

Reverses if: sovereign funding lapses or open-lab economics fail to scale.

Trust & safety

Under active tracking: misuse capability and how easily safety tuning strips from open weights, hard-friction zones (above all synthetic CSAM and NCII), and whether NTIA's monitoring posture holds.

Reverses if: a major misuse event, or a shift from monitoring to restriction.

There is a test you can run for the rest of this. Look at who is seated in the rooms where AI gets decided, and with what status. The day they seat the people who keep AI open, portable, and widely deployed on equal footing, the shift from renting to owning will have happened. The window is open now. It is closing slowly enough to be easy to ignore, and the lease is shorter than it looks. Build with us.

This is v1.0.1. We'd like to hear from you.


Citations

Section 1 · The current state of open-source AI

Section 2 · Who's betting on it

Section 3 · Why it's happening everywhere

Section 4 · The harness is the new frontier