TJF (@TJWXF3) on X

X (formerly Twitter) ·

17 min read Original article ↗

A firm cannot own its AI learning loop if it cannot verify the trace.

The most useful framing of the AI transition this week was not about model benchmarks. It was about capital.

In a recent post, Satya Nadella framed the future of the firm around two forms of capital: human capital, meaning judgment, relationships, pattern recognition, and taste, and token capital, the AI capability a company builds and owns.

A model is something you rent. A workflow is something you operate. But capital is something that compounds.

The real question is whether companies will actually own that compounding loop. A firm can use agents all day and still own nothing durable if the traces of those actions live inside a vendor log, disappear with a model migration, or cannot be verified when someone with authority asks what happened.

A learning loop without receipts is just a faster way to accumulate unverifiable actions.

The capital is in the trace, not the model.

Trust did not disappear. It moved.

There is a familiar shape to this problem.

Almost ten years ago, I wrote about cryptoeconomics as a way to build digital systems we could trust. The point was not that blockchains magically remove trust. They do not. The point was that trust can be relocated: away from opaque intermediaries and into cryptography, incentives, protocol rules, and open verification.

The AI-agent transition brings back the same problem in a new form. The question is no longer only whether we can trust a ledger, a platform, or a counterparty. It is whether a firm can trust a digital actor that reads, reasons, writes, calls tools, and proposes actions on its behalf.

The answer rhymes with the old one: do not ask the actor to be trusted. Change the system so the action can be verified.

Trust did not disappear. It moved again, this time to the action boundary.

There is a longer arc here too. Writing, print, the internet, and real-time networks each widened the circle of people who could act on shared information. But every one of those substrates coordinated humans communicating with humans. Agents are different. They do not just relay information; they take actions in the world.

Coordinating them needs a substrate for verifiable action, not just shared speech: a way to know, across firms and counterparties and time, what an agent was authorized to do and what it actually did.

Without it, agentic activity simply outruns our ability to account for it.

Figure 1. The trust boundary moved from platform, to protocol, to agentic action. Agentic systems need proof where intention becomes action.

Capital is not a tool

Nadella's frame matters because it separates capital from tools.

A tool is bought, used, and replaced. It may make a worker faster, but it does not necessarily compound inside the firm. Most AI products today are tools: agents that draft emails, write code, run searches, summarize documents, and generate reports.

Useful, but not automatically capital.

Token capital is different. It is the owned capability that gets stronger with use because the firm retains the traces, evals, policies, context, approvals, and outcomes that make improvement possible.

Ownership here is not just legal. It is operational. A firm has to be able to trace what the system did and why, verify that trace independently of any vendor, switch the underlying model without losing accumulated expertise, evaluate agents against outcomes that matter to the business, and disclose selectively, proving control without exposing trade secrets.

That is a proof problem, not a model problem.

A model can generate an answer. A receipt can preserve the institutional fact that the answer was proposed, evaluated, approved, denied, or used.

That difference is the line between rented tooling and owned capital.

Receipts are the memory layer

The learning loop everyone wants looks like this:

human judgment to agent action to outcome to private eval to improved workflow to better agent action

But for that loop to compound, each step must produce a durable, inspectable record.

Not a screenshot. Not a Slack message. Not a database row that disappears when a vendor contract ends. Not a log file buried inside the same system whose behavior is being reviewed.

A receipt is a signed record created at the moment of decision. It binds some combination of: agent identity; human owner; delegated authority; trusted context; policy version; proposed action; approval or denial; outcome reference; verification key; timestamp.

The point is not to record everything forever. The point is to preserve the decision trace in a form the firm owns and can verify.

A receipt is institutional compression: it turns a messy moment of authority, context, policy, action, approval, and outcome into a portable object the firm can carry forward.

That gives the learning loop three properties it otherwise lacks.

Portability across models. The receipt is tied to the action, not to the weights that produced it. A firm can use one frontier model today, a different one tomorrow, and a local or fine-tuned model later, while keeping institutional memory intact.

Auditability without exposure. A receipt can prove that a specific action was evaluated under a specific mandate without revealing the underlying position, prompt, source document, or proprietary context. In financial, legal, healthcare, and enterprise settings, that distinction is the whole game: the evidence must be strong enough to verify and narrow enough not to leak the edge.

Human accountability. Token capital grows through human direction, but direction only counts if it is attributable. A receipt binds a human approval to the exact action, context, and policy state being approved. That is the institutional version of this is what I authorized.

What receipts prove, and what they do not

A receipt is not magic.

It does not prove the model was right, the trade was good, the document was wise, or the firm's judgment was sound.

It proves something narrower and more valuable: what action was proposed, what authority the agent had, what context was used, what policy was applied, whether the result was allow, deny, escalate, or co-sign, who approved it, and whether the record was later altered.

That narrower proof is what makes the broader learning loop possible.

Without it, private evals are built on sand. You cannot reliably score outcomes if you cannot reconstruct the decision. You cannot switch models if the traces are trapped in a vendor system. You cannot prove mandate adherence if the evidence requires disclosing the very thing the mandate is meant to protect.

The receipt is not the eval. It is the evidentiary substrate that makes private evals meaningful.

The loop is being described from three directions at once

Here is what makes this moment strange: the same loop is being described from three directions at once, and none of the three accounts can verify its own trace.

Nadella names it as capital: the firm's owned, compounding AI capability, the thing that must survive a model swap.

Jack Dorsey and Roelof Botha push it into org structure: a company world model built from machine-readable artifacts that carries what middle management used to carry, coordinating work without humans relaying information up and down a chain.

Andrej Karpathy takes it down to the individual: a persistent, interlinked knowledge base that the model compiles and maintains, so judgment accumulates instead of being re-derived on every query.

Three scales: enterprise, organization, person.

One shared shape: human judgment compounding into a firm-owned learning loop.

And one shared blind spot.

Nadella's capital is only capital if you can prove you own it. Dorsey and Botha's world model creates exactly the trace this essay is concerned with: powerful, operationally useful, and potentially unverifiable unless the authority boundary is signed.

When the intelligence layer surfaces a financial offer or routes resources to a customer, what record proves it acted within mandate, on trusted context? When a DRI has "full authority to pull resources" across teams, what is the signed record of that authority, with no manager left in the chain to later say I authorized that?

And Karpathy's wiki is honest about its own mechanics: the raw sources stay immutable, but the synthesis layer is built by the model, rewritten by the model, its contradictions flagged by the model, a self-asserted artifact with no independent verifier and no proof of which source justified which claim.

Each is a learning loop without receipts.

The logic compounds in the worst direction: the more coordination, synthesis, and authority you move into the system, the more load-bearing the proof at the boundary becomes, because there is no longer a human in the chain who can be asked what they authorized.

Strip out the human relay and you strip out the human who could be questioned.

The proof has to live at the action boundary instead.

That is the negative space these three visions leave open. It is the space the receipt layer is built to fill.

The circular trust problem, and why the verifier must be blind

Which raises the obvious objection.

If the answer to "trust the vendor's logs" is "trust a receipt layer instead," haven't we just moved the circular-trust problem one box to the left?

If the agent is built by a vendor, the model served by that vendor, the tools mediated by that vendor, and the audit trail stored by that vendor, the firm is asking the system under review to grade its own behavior.

That may be fine for low-risk productivity software. It is not enough for workflows where the evidence may be shown to a board, regulator, insurer, or counterparty.

But a receipt layer that you have to call me to verify has exactly the same defect.

If verification requires an API call to the party whose conduct is being verified, it is not institutional proof. It is a hosted assertion.

So the receipt layer has to be structurally independent, not just of the model vendor, but of the receipt issuer too.

This is the part that has to be mechanical, not rhetorical.

At ScopeBlind, the receipts are issued under Veritas Acta, an open protocol with an issuer-blind verifier.

Three properties make the independence real:

The verifier is open source. Anyone, a board, an auditor, a regulator, a counterparty, can run it themselves. Verification is not a service I sell; it is code you execute.

The issuer is blind. The cryptography is structured so that ScopeBlind, as issuer, does not see the contents of what is being verified. There is nothing to leak and no privileged position to abuse, because the issuer never holds the underlying secret.

Verification runs offline. Checking a receipt requires no live call to ScopeBlind and no call to the model vendor. The proof stands on its own verification path. It survives model migration, vendor termination, pricing changes, outages, and disputes, including a dispute with ScopeBlind.

This is the difference between "trust our logs" and "here is a proof you can check without trusting anyone, including us."

Selective disclosure does the rest: prove the control without revealing the secret.

That is the only version of a receipt layer that actually escapes the circular-trust trap rather than relocating it.

The liability shift

When a human breaches a mandate, the organization has a familiar control story: who acted, what authority they had, what policy applied, who supervised them, what record exists.

Agentic systems blur that story.

If an agent accesses the wrong file, proposes an out-of-mandate action, discloses restricted information, or acts on corrupted context, the firm cannot answer with the model did it. The obligation still sits with the organization that deployed the system.

That is why receipts matter to risk officers, not just engineers.

You cannot govern, insure, or defend a workflow you cannot reconstruct.

A receipt does not eliminate liability. It changes the evidentiary posture. The firm can show what authority existed, what context was used, what policy was applied, whether human co-sign was required, and whether the action was allowed, denied, or blocked before execution.

That is the difference between we think the control existed and here is the proof boundary.

There is also a timing problem.

Agentic systems can be deployed in weeks, but boards, insurers, auditors, allocators, and regulators often arrive months or years later. By then, the model has changed, prompts have changed, employees have left, vendors have migrated logs, and the workflow has been rewritten.

That is the regulatory clock.

The question will not be what the agent can explain today. It will be what the firm can prove later.

Receipts move evidence creation from the future investigation back to the moment of action.

Negative-space proof

The most valuable record is not always a record of what happened.

In agentic systems, the more important proof is often what did not happen: the file the agent could not read, the tool it could not call, the order it could not stage, the report it could not export, the action it could not take without human co-sign.

Most observability systems log the positive space, the path taken. A receipt layer should also prove the negative space: that a fail-closed gate evaluated the request and blocked, constrained, or escalated it under the active mandate and context.

That is the difference between surveillance and restraint.

Surveillance tells you what the agent did.

Restraint proves where the agent stopped.

This proof is scoped. It proves restraint inside the governed action path. Enterprise controls still have to make that path mandatory for sensitive actions.

Figure 2. Logs show the path taken. Restraint receipts prove where the boundary held.

Source context is part of the proof boundary

A receipt binds an action to the context used at decision time. It does not automatically prove that the source file was complete, authentic, current, or economically correct.

That is why the context layer matters: source lineage, file hashes, freshness checks, parser confidence, independent feeds, and human approval of low-confidence parses all become part of the proof boundary.

A firm should not merely ask what did the agent decide? It should ask: what context was the agent allowed to use, who approved it, how fresh was it, what parsed successfully, what failed, and what did the gate do when confidence was low?

The dangerous failure mode is not only a hallucinating agent. It is an agent acting confidently on corrupted context.

Make it concrete.

An agent pulls a counterparty exposure file to size a position. The file is a day stale because an overnight feed silently failed, and a parser quietly dropped two rows it could not read. The agent does exactly what it was told, on the data it was given, and proposes a sensible-looking trade.

It clears.

Three months later, a regulator asks why the limit was breached.

The logs show the agent ran and the trade went through. They do not show that the context was stale, that two rows were missing, or that no one approved acting on a low-confidence parse, because nothing recorded the state of the inputs at decision time.

The firm cannot prove the control existed, because operationally it did not.

A receipt would have bound the action to that exact context: feed timestamp, parser confidence, the rows that failed, and the gate's decision to proceed or escalate.

With it, the firm can show precisely what was known and what was not.

A serious receipt layer must fail closed when the context is stale, ambiguous, incomplete, or outside the approved source scope, and prove that it did.

Receipts as private-eval substrate

Private evals are only as good as the traces they score.

If a trace does not bind context, authority, policy, approval, action, and outcome, the eval cannot distinguish a good model from a lucky outcome, a bad model from corrupted context, a policy failure from a prediction failure, a human override from an agent error, or a workflow improvement from a shift in the external environment.

That is why receipts are not just compliance artifacts.

They are training-signal governance.

A firm-owned learning loop needs firm-owned traces, otherwise the most valuable part of the system, the accumulated record of judgment, context, action, and outcome, lives outside the firm's control.

No company should spend years turning its people's judgment into someone else's model memory.

Figure 3. Human capital becomes token capital only when the firm owns the trace.

The anti-commoditization case

Nadella is right to warn against a future where a small number of AI systems capture all the economic returns. That outcome is not only unjust. It is unstable. Firms will not tolerate becoming tenants of their own expertise.

There is a structural choice buried in this.

One reading of the last decade casts two technologies as opposing forces: decentralizing tools that push power outward, and centralizing intelligence that concentrates it. Framed that way, an AI future looks like a one-way pull toward concentration, value accruing to whoever owns the largest model.

But that dichotomy is not fixed.

Customer-owned token capital, made portable and verifiable by receipts, is precisely the thing that lets the decentralizing impulse survive inside a world of powerful central models.

The receipt layer is how a firm keeps its compounding expertise its own even while renting frontier intelligence from someone else.

It is not decentralization versus AI.

It is the mechanism that lets ownership and frontier capability coexist instead of one eating the other.

The sovereignty problem is not solved by using a better model. It is solved by owning the loop around the model.

When a company's workflows, domain knowledge, and accumulated judgment are encoded as signed traces the company controls, value accrues to the company rather than only to the model provider.

The model becomes interchangeable infrastructure.

The receipt layer becomes part of the firm's owned, compounding asset base.

This is why compliance is a cost and proof is an asset.

Compliance is paying someone to reconstruct what happened after the fact.

Proof is producing the evidence at the moment of action, in a format that outlives every vendor.

What we are building

At ScopeBlind, we are building receipt infrastructure for agentic systems, on the open Veritas Acta protocol.

The first wedge is financial workflows, where agents may soon help with research, trade proposals, risk review, reporting, and operational processes, but cannot be given unchecked authority.

To make the loop ownable and provable, the receipt layer has to do six jobs:

Signed authority: what the agent is allowed to access, call, propose, or disclose.

Trusted context: the files, data, mandate, or book state the action was evaluated against.

Fail-closed gates: deterministic policy decisions before action, not audit after.

Human co-sign: approvals bound to the exact payload and context.

Restraint receipts: proof of what was blocked or escalated, not only what was allowed.

Open, issuer-blind verification: receipts checked offline, without a hosted ScopeBlind service or a model-vendor log, and selective disclosure that proves the control without revealing the secret.

The goal is not to make agents autonomous.

The goal is to make useful agent workflows permissible.

A receipt does not prove the agent was right.

It proves the control was real.

The primitives here are not proprietary magic. The signing, offline verification, hash-chaining, and selective-disclosure patterns are standard cryptography, deliberately so, because proof that depends on a vendor's secret sauce is not proof.

I wrote the receipts lesson in Microsoft's open AI Agents for Beginners course, which walks through producing and verifying an Ed25519-signed receipt in about fifty lines of Python, and the same care goes into the line between what a receipt proves (attribution, integrity, ordering) and what it does not: that the action was correct, or that the policy was actually enforced.

The Veritas Acta protocol formalizes that format. The verifier is open so you never have to take my word for any of it.

The test of sovereignty

Nadella's test is the right one: a company should be able to switch out a generalist model without losing the "company veteran" expertise built into its learning system.

That test is only passable if the expertise is stored somewhere the company owns, in a format that outlasts the model.

Receipts are that format.

A frontier model is impressive. A frontier ecosystem is stable. But a frontier ecosystem without proof is just a market of faster black boxes.

The stable equilibrium is one where every organization owns the learning loop that encodes its institutional knowledge, and can prove it.

Trust does not disappear. It moves.

The capital is in the trace, not the model.

A model is a tool you rent. A receipt is capital you own.

If you are building agentic systems and want the learning loop to compound inside your firm, get in touch, read the receipts lesson in Microsoft's AI Agents for Beginners, or look at the Veritas Acta draft.

Notes and sources

Satya Nadella, on human capital and token capital. https://x.com/satyanadella/article/2066182223213293753

Jack Dorsey and Roelof Botha, on Block as an intelligence rather than a hierarchy. https://block.xyz/inside/from-hierarchy-to-intelligence

Andrej Karpathy, on the LLM-maintained personal wiki pattern. https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f

Tom Farley, "Cryptoeconomics We Can Trust" (c. 2016).

Tom Farley, "Securing AI Agents with Cryptographic Receipts," Lesson 18, Microsoft AI Agents for Beginners.

IETF Internet-Draft, "Signed Decision Receipts for Machine-to-Machine Access Control" (draft-farley-acta-signed-receipts).

Veritas Acta protocol and open verifier. https://github.com/VeritasActa/verify