The Thinking Wattage of AI Agents vs. the Human Brain

7 min read Original article ↗

AI inference efficiency keeps pushing on. A few years ago the wisdom was that the human brain is WAY beyond the efficiency of GPUs in data centers. But with the OpenAI Jalapeño chips and all the other advancements on one side, and the fact that, hmm, I don't actually type or read infinitely fast, I started thinking — are we closer to parity at least on "run of the mill" type of thinking — the things that one does in a workday that are not feats of genius or accumulated experience condensing into a great decision? Even my laptop already types code faster than me, even if Qwen 3.8 27B might feel painfully slow.

So I asked Opus 5 to do the arithmetic properly, with the numbers from Vera Rubin and OpenAI's Jalapeño on the table. The answer below is published as it came out, unedited. Here's what I asked:

With the latest AI models and Nvidia's Vera Rubin platform token efficiency or the new OpenAI Jalopeno chip, what is the effective "thinking wattage" of AI agents today vs. human brain. A brain is a lot more economic, but on the other hand it is also only providing "office thinking" with some of that capacity. Build me a range based on latest data and educated guesses where the data is scarce (e.g. frontier models — maybe we consider the near-frontier Chinese open weights models unless we have good guesses for Sol or Opus). Extrapolating forward, are we months, years or decades from parity, or already past some limit?

The answer

Short version: you have to split "thinking wattage" into three different quantities, because they give answers that differ by four orders of magnitude — and two of the three crossovers have already happened.

The AI side, grounded

Closed frontier models (Opus-class, Gemini-class, GPT-class) publish nothing, so I'll use your suggestion and anchor on near-frontier open weights, which are now genuinely close in capability: Kimi K3 at 2.8T total / 104B active, DeepSeek-V4-Pro at 1.6T / 49B active, GLM-5.1/5.2 at ~744B / 40B active. Frontier closed models are probably in the same 50–150B-active band, so the arithmetic transfers.

Measured energy per output token, production-representative serving:

Source Configuration J / output token
ML.Energy (B200, GPU only) Qwen3 235B-A22B Thinking FP8, 4×B200, min-energy config: ~0.4 J/token 0.4
ML.Energy Qwen3 32B dense, chat vs reasoning: 0.15 → 0.31 J/token 0.15–0.31
Microsoft/Oviedo (Joule, 2026) 0.31 Wh median per frontier query (>200B, H100), 3.91 Wh for 5,000-token reasoning ~2.8–3.7

The ML.Energy numbers are GPU-only; accelerators are 50–70% of datacenter power, so multiply by ~1.5–2× for wall energy. That gives a defensible 2026 band of ~0.6–4 J per output token at the wall for a frontier-class model, with ~1 J/token as the best-case well-batched production figure and 10+ J/token for tight-latency, long-context, or low-concurrency serving.

Both new platforms move this, but less than the headlines suggest. Jalapeño is 13.4 PFLOP/s of mxfp4 with 15.4 TB/s across 216 GiB HBM4 in a 700 W package, and its Hot Chips numbers are ~1.5–1.9× peak tokens/s/kW versus GB200/GB300, measured on the SemiAnalysis InferenceX Pareto frontier with TDP normalisation. Nvidia claims 8× inference perf/watt over Blackwell for Vera Rubin NVL72 (3.6 EFLOPS FP4 inference per rack) — but that's a marketing stack including MTP, NVFP4 and rack-level accounting, and SemiAnalysis found Jalapeño and Vera Rubin essentially tied on TCO per token, with Rubin's July figures inflated by multi-token prediction. Realistically call it 2–4× real improvement in J/token over B200-class by 2027, so ~0.2–1 J/token.

Three wattages

(1) Power while thinking at human pace. Human inner speech / reading runs ~250 wpm ≈ 5 tokens/s. At 0.6–4 J/token that's 3–20 W. Brain is ~20 W. Already at parity, and probably past it. This crossover happened quietly around 2024–2025 and nobody noticed because it's the least economically meaningful of the three.

(2) Raw operations per joule. Rubin NVL72: 3.6 EFLOPS FP4 at ~125 kW ≈ 2.9×10¹³ FLOP/J. Jalapeño: 1.9×10¹³ FLOP/J at package level. The brain, at 10¹⁴–10¹⁵ synaptic events/s on 20 W, is 5×10¹²–5×10¹³ ops/J. Silicon is now within a factor of ~2–10 of biological synaptic efficiency, and by some counts ahead. Hardware is no longer the gap.

(3) Power per unit of delivered work. This is the one that matters, and it's where AI still loses. A frontier agentic session — reasoning traces, tool calls, retries, verification, prefill on long context — burns 1–20M tokens for a day's work. Reasoning alone costs 25× more energy per response than chat: 10× more output tokens plus higher J/token from KV-cache pressure limiting batch size.

Agent workload @1 J/tok, over 8h over 24h
1M tok/day (light) 35 W 12 W
5M tok/day (typical coding agent) 174 W 58 W
20M tok/day (heavy multi-agent) 694 W 232 W

So ~30–700 W of effective "office thinking wattage" versus a 20 W brain: currently 2–35× worse.

The human side, honestly accounted

Your instinct about "office thinking" being a fraction of the 20 W is right, but it cuts the other way from how it's usually framed. Task-evoked metabolic increase in the brain is only a few percent — the marginal cost of hard thinking is well under 1 W. The 20 W is mostly overhead: keeping 10¹⁴ synapses alive, vision, motor, homeostasis. You can't buy the 1 W without the 20 W.

And the 20 W isn't the real accounting boundary either:

  • Body: ~100 W continuous, ~2.4 kWh/day.
  • Childhood amortisation: 20 years at ~100 W = 17.5 MWh metabolic, spread over ~80,000 working hours = +220 W per working hour. (The AI equivalent — a 10²⁶–5×10²⁶ FLOP training run at 30–300 GWh, amortised over 10¹⁵–10¹⁶ served tokens — is only 1–100 mJ/token, i.e. 1–10% of inference. Training amortises far better than childhood does.)
  • Societal: Finland runs ~7.5 kW of primary energy per capita continuously; EU electricity alone is ~0.7 kW/capita. A human knowledge worker's fully loaded energy footprint is kilowatts, not watts.

There's also the Zheng & Meister result worth putting on the table: human information throughput maxes out around 10 bits/s despite sensory input at ~10⁹ bits/s. At 20 W that's 0.5 bits/J of conscious cognition. It's a behavioural bottleneck rather than a computational limit, but it means the brain's deliverable bandwidth is tiny relative to its substrate — which is precisely why it can be beaten on (1) while still winning on (3).

Are we months, years, or decades out?

  • Already past: per-token wattage at human speech rate (crossed ~2024–25), and raw ops/joule versus synaptic ops/joule (crossing now).
  • Years, not decades, for the meaningful one. Hardware efficiency is improving ~40% per year; model efficiency is faster — token price half-lives of 1.10 years (economy tier) and 1.55 years (mid tier) plus the ongoing shift to sparse MoE, where a 30B-A3B model uses 3.56× less energy per token than a dense 32B. Net ~3–10× per year in joules per unit of capability, against a countervailing ~1.5–3×/year growth in tokens-per-task from agentic scaffolding. Call it 2–5× net annual improvement, closing a 2–35× gap: crossover on energy-per-day-of-office-work between 2027 and 2030.
  • Already past, if you use the right boundary. Against a human's fully loaded ~1–7 kW societal footprint rather than the brain's 20 W, agents at 30–700 W are already the cheaper way to buy a unit of cognition. This is the number that will actually drive procurement decisions, and it's why the buildout continues despite the per-brain comparison looking unflattering.

Where the remaining gap actually lives. Not in hardware. At 2×10¹³ FLOP/J and 2×(49–104B) active params, the arithmetic roofline is 5–30 mJ/token — but real serving delivers ~0.6–4 J/token. That's a ~100–300× utilisation gap, caused by memory-bandwidth-bound decode, batch limits from KV-cache pressure, and idle silicon between phases. This is exactly what OpenAI is attacking: Jalapeño keeps KV local and varies the active silicon mix per phase so unused units go dark, rather than paying baseline package/HBM/network power on a heterogeneous fleet. Two orders of magnitude of headroom exist without any new physics — Landauer is still ~10¹⁹× away and irrelevant to this timescale.

The honest bottom line: the brain isn't winning on thermodynamics anymore. It's winning on sample and step efficiency — the number of "thoughts" needed to finish a task. That gap is algorithmic, and it's the one worth watching rather than the watts.