TL;DR
- Every frontier model passes the pelican-on-a-bicycle test now. It is in every training corpus and was never precisely gradable, so it stopped separating anyone.
- Our replacement: draw a MacBook Pro 16 in SVG, one shot. Everyone knows the exact proportions, so a 5% error is visible instantly.
- Claude Fable 5 draws the best one. Gemini 3.6 Flash gets startlingly close for $0.16 - a thirteenth of Fable's best run.
- New: Claude Opus 5 (Anthropic) draws the most detailed Anthropic MacBook here - but the geometry is broken at every effort level, it only looks the part, and thinking-on-by-default burns 20-47K tokens per drawing (10-25x Opus 4.8's cost); its
maxrun spends the whole budget thinking and never draws the laptop. - The
maxeffort setting is mostly a trap: it buys cost and latency, sometimes a run that never terminates, and never an understanding of the object. - 64 runs across 12 models, both tasks, every effort level. Unedited output, providers' own billing.
NEW2026-07-24: Added Claude Opus 5 (Anthropic) at high/xhigh/max, both tasks - the most detailed Anthropic laptop here, but the geometry is broken at every level (skewed at high, broken isometry at xhigh, no laptop at all at max), and thinking-on-by-default makes it 10-25x pricier per drawing than Opus 4.8.
For two years the best quick test of a new model was Simon Willison's pelican: ask for a pelican riding a bicycle as SVG and look at the result. It worked because it could not be faked. It no longer separates anyone - so we built the test we actually wanted: draw a MacBook Pro 16 in SVG, in a dark colour, one shot. Everyone reading this knows exactly what a MacBook looks like, so everyone reading this can grade it. Below is every run we have: 12 models, both tasks, the full effort ladder, 64 runs, unedited.

Why the pelican had to go #
A benchmark is useful exactly as long as it separates models. We ran the pelican as a control through the identical harness - same models, same effort ladder, same one-shot rule - and every single model produced a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading. Two problems got it there:
- Everyone has trained on it. It is the most famous drawing prompt in AI. Three years of pelicans and public commentary about them sit in every training corpus. A good pelican no longer tells you whether the model understands geometry or has simply seen ten thousand graded pelicans.
- It never had a gradation. Nobody knows how long a pelican's beak should be relative to its wingspan, so grading tops out at "recognisable". There is no scale on which one passing pelican beats another.
A MacBook Pro fixes both. Hundreds of millions of people have stared at one for thousands of hours, so the lid, the keyboard grid, the notch, the hinge and the trackpad have a precise shared ground truth - your eye flags a small error instantly. And it has a real difficulty ladder: dozens of parts plus a 3D perspective view. Models can now fail at many different heights instead of clustering at "recognisable". The numbers say the same thing: the models spend roughly 4x fewer output tokens on a pelican than on the MacBook (Opus 4.6: 2,091 vs 8,156), and cranking effort barely moves the pelican at all.
How we ran it #
Create an SVG 3D model of a MacBook Pro 16 in a dark color. Respond with ONLY the SVG markup, no explanation.
One API call per cell. No retries, no system prompt, no examples, no cherry-picking - the first answer is the answer. Each model runs the effort ladder its provider supports (high, xhigh, max); Grok 4.5 tops out at high, Gemini 3.6 Flash exposes low/medium instead of xhigh, and Kimi K3 runs low/high/max with no xhigh. Costs are the providers' own usage-reported billing. Runs marked * were requested as xhigh and clamped to high by the model's generation.
On the images. Each drawing is a browser render of the model's unedited SVG - nothing touched up, and the raw SVGs stay in the repo. We rasterise deliberately: the models all reach for the same generic ids (#screenGrad, #screenClip, #shadow), so inlining fifty of them into one page makes the browser resolve one model's clip-path into another model's drawing. Rendering each SVG in its own document is the only way to show you what the model actually drew.
On reproducibility. We re-ran the pelican ladder nine days later on a different output budget; token counts landed within about 10% of the originals (Fable at high: 2,264 then 2,002). Each run below records the sweep it came from, so mixed dates stay visible rather than quietly blended.
Every run #
The whole matrix, both tasks. Filter to one task or vendor, sort by cost to line them up cheapest-to-priciest, or show only the newest model.
64 of 64 runs

$0.50 · 20,145 tokens · 4.3 min
Anthropic · run 2026-07-24
All the detail is here - notch, dock with app icons, a code window, "MacBook Pro" under the screen - but the base is skewed: the deck rakes and runs too deep, in a slightly different perspective from the lid. $0.50 and 20K output tokens, because thinking is on by default.

$1.18 · 47,237 tokens · 8.5 min
Anthropic · run 2026-07-24
More detail and better materials than high, but the isometry is broken: the screen and the keyboard deck sit in different perspectives, so it looks slick at a glance and falls apart on a second look. 47K output tokens and $1.18 for it.
Runaway. Thinking is on by default, so it spent almost the entire 64K budget reasoning and emitted only the background glow before the cap - no laptop at all, $1.60 billed for nothing. The same failure mode as Sonnet 5.
$1.60 · 64,000 tokens · 13.3 min
Anthropic · run 2026-07-24

$0.06 · 2,290 tokens · 28s
Anthropic · run 2026-07-24
Clean and correct first try - white body, big orange pouch, both wheels properly spoked, legs on the pedals.

$0.09 · 3,681 tokens · 46s
Anthropic · run 2026-07-24

$0.15 · 5,957 tokens · 76s
Anthropic · run 2026-07-24
The dial that does nothing on other pelicans does something here: max adds a dynamic pose, speed lines and a curved handlebar over high. On a saturated task, that is polish, not correctness.

$0.15 · 9,846 tokens · 5.0 min
Moonshot AI · run 2026-07-24
A proper dark MacBook - dock, notch, keyboard, trackpad - already here for $0.15, a hair over Gemini 3.6 Flash at max.

$0.46 · 30,579 tokens · 17.5 min
Moonshot AI · run 2026-07-24
Max earns it here - one of the best laptops in the benchmark: correct 16-inch proportions, speaker grilles, a code window, no visible errors. Fable-tier for a fifth of the price, but 30,579 tokens and 17.5 minutes.

$0.05 · 3,425 tokens · 108s
Moonshot AI · run 2026-07-24
Aces it first try - clean bike, big orange pouch, a crest and motion lines.

$0.09 · 6,183 tokens · 3.5 min
Moonshot AI · run 2026-07-24
6,183 tokens against high's 3,425 for a pelican that is no better - the effort dial doing nothing again.

$0.07 · 9,544 tokens · 36s
Google · run 2026-07-21

$0.08 · 11,317 tokens · 47s
Google · run 2026-07-21

$0.16 · 21,051 tokens · 83s
Google · run 2026-07-21
Max effort, complete: a detailed dark laptop in a proper three-quarter view, for $0.16.

$0.10 · 12,996 tokens · 54s
Google · run 2026-07-21

$0.09 · 12,340 tokens · 50s
Google · run 2026-07-21

$0.09 · 11,520 tokens · 46s
Google · run 2026-07-21

$0.27 · 5,383 tokens · 68s
Anthropic · run 2026-07-13

$0.73 · 14,554 tokens · 3.0 min
Anthropic · run 2026-07-13
Winner of the launch sweep - the most detailed MacBook anyone drew inside the 16K cap.

$2.13 · 42,560 tokens · 9.8 min
Anthropic · run 2026-07-21
Uncapped (64K, streamed): the best drawing in the benchmark - isometric, dock, side ports - for $2.13 and 9.8 minutes.

$0.004 · 668 tokens · 36s
xAI · run 2026-07-13
high is its ceiling. 668 output tokens - it barely tried.

$0.11 · 2,264 tokens · 26s
Anthropic · run 2026-07-13

$0.12 · 2,400 tokens · 30s
Anthropic · run 2026-07-13

$0.43 · 8,557 tokens · 112s
Anthropic · run 2026-07-13

$0.009 · 1,460 tokens · 17s
xAI · run 2026-07-13

$0.26 · 8,551 tokens · 109s
OpenAI · run 2026-07-13

$0.45 · 15,042 tokens · 3.5 min
OpenAI · run 2026-07-13
Timed out. It did not return in 36 minutes, even with a 64K budget - exactly as in the first run.
- · - tokens · 36.6 min
OpenAI · run 2026-07-21

$0.09 · 5,685 tokens · 34s
OpenAI · run 2026-07-13

$0.17 · 11,650 tokens · 104s
OpenAI · run 2026-07-13

$0.45 · 29,834 tokens · 4.6 min
OpenAI · run 2026-07-21
Uncapped it finishes - and the form is still wrong. Budget does not buy understanding.

$0.09 · 3,088 tokens · 42s
OpenAI · run 2026-07-13

$0.11 · 3,668 tokens · 51s
OpenAI · run 2026-07-13

$0.44 · 14,690 tokens · 4.0 min
OpenAI · run 2026-07-13
Sol does finish at max here. The task it cannot finish is the MacBook.

$0.05 · 3,083 tokens · 23s
OpenAI · run 2026-07-13

$0.10 · 6,445 tokens · 54s
OpenAI · run 2026-07-13
Truncated at the 16K cap.
$0.24 · 16,000 tokens · 2.5 min
OpenAI · run 2026-07-13

$0.03 · 2,644 tokens · 29s
Anthropic · run 2026-07-13

$0.12 · 12,011 tokens · 2.0 min
Anthropic · run 2026-07-13
A genuine runaway: it burned the entire 64K budget and still never closed the SVG.
$0.66 · 65,538 tokens · 11.6 min
Anthropic · run 2026-07-21

$0.01 · 1,179 tokens · 10s
Anthropic · run 2026-07-13

$0.02 · 1,769 tokens · 14s
Anthropic · run 2026-07-13
Truncated: it hit the 16K output cap without ever closing the SVG.
$0.16 · 16,000 tokens · 3.0 min
Anthropic · run 2026-07-13

$0.20 · 8,156 tokens · 85s
Anthropic · run 2026-07-13

$0.21 · 8,328 tokens · 82s
Anthropic · run 2026-07-13

$0.19 · 7,767 tokens · 76s
Anthropic · run 2026-07-13

$0.05 · 2,091 tokens · 32s
Anthropic · run 2026-07-13

$0.06 · 2,461 tokens · 33s
Anthropic · run 2026-07-13

$0.05 · 1,942 tokens · 27s
Anthropic · run 2026-07-13

$0.05 · 2,087 tokens · 21s
Anthropic · run 2026-07-13

$0.10 · 3,912 tokens · 39s
Anthropic · run 2026-07-13

$0.18 · 7,189 tokens · 65s
Anthropic · run 2026-07-13

$0.03 · 1,083 tokens · 14s
Anthropic · run 2026-07-13

$0.03 · 1,183 tokens · 14s
Anthropic · run 2026-07-13

$0.04 · 1,528 tokens · 16s
Anthropic · run 2026-07-13

$0.11 · 7,151 tokens · 82s
Anthropic · run 2026-07-13

$0.11 · 7,340 tokens · 85s
Anthropic · run 2026-07-13

$0.10 · 6,809 tokens · 77s
Anthropic · run 2026-07-13

$0.03 · 1,999 tokens · 25s
Anthropic · run 2026-07-13

$0.03 · 1,854 tokens · 23s
Anthropic · run 2026-07-13

$0.03 · 2,000 tokens · 24s
Anthropic · run 2026-07-13

$0.007 · 2,348 tokens · 14s
Google · run 2026-07-13

$0.007 · 2,328 tokens · 13s
Google · run 2026-07-13

$0.006 · 1,897 tokens · 12s
Google · run 2026-07-21
Only possible after we wired Gemini to thinking_level - this model had no working effort dial in the launch sweep.

$0.008 · 2,536 tokens · 14s
Google · run 2026-07-21
Model by model #
Each model with both of its drawings side by side - the pelican it aces and the MacBook that actually grades it. Newest first.
Claude Opus 5 NEW#
Anthropic · in Playcode since 2026-07-24
The most detailed, ambitious Anthropic laptop here - notch, dock, traffic lights, a code window, convincing materials - and the first Anthropic model whose effort dial actually changes the drawing instead of redrawing the same thing. But it fails the one thing the MacBook test is for: the geometry. At high the base is skewed; at xhigh the isometry is broken outright - the screen and the keyboard deck sit in different perspectives, so it reads slick at a glance and falls apart on a second look. And max is the worst of the three: thinking is on by default, so it spent almost the entire 64K budget reasoning and emitted only the background glow before the cap - no laptop at all, $1.60 for nothing. A clear step up from Opus 4.8's flat mid laptop in detail and effort-response - but what separates Fable is that it makes no visual errors, where Opus 5 only looks the part.
MacBook

$0.50 · 20,145 tokens · 4.3 min

$1.18 · 47,237 tokens · 8.5 min
Runaway. Thinking is on by default, so it spent almost the entire 64K budget reasoning and emitted only the background glow before the cap - no laptop at all, $1.60 billed for nothing. The same failure mode as Sonnet 5.
$1.60 · 64,000 tokens · 13.3 min
Pelican

$0.06 · 2,290 tokens · 28s

$0.09 · 3,681 tokens · 46s

$0.15 · 5,957 tokens · 76s
Kimi K3 #
Moonshot AI · in Playcode since 2026-07-23
The new challenger, and a serious one. It draws a proper dark MacBook - dock, notch, speaker grilles - at both levels, and at max it produces one of the best laptops in the benchmark: a clean render with correct 16-inch proportions and no visible errors, for $0.46. That is a fifth of Fable's price for a drawing in the same class. The catch is the clock - that run took 17.5 minutes - and, as ever, the pelican it already aces at high gains nothing from max.
MacBook

$0.15 · 9,846 tokens · 5.0 min

$0.46 · 30,579 tokens · 17.5 min
Pelican

$0.05 · 3,425 tokens · 108s

$0.09 · 6,183 tokens · 3.5 min
Gemini 3.6 Flash #
Google · in Playcode since 2026-07-21
The surprise of the whole benchmark. Google's newest Flash is the first Gemini here with a working effort dial, and at max it draws a near-perfect 3D MacBook for $0.16 - a fifth of Fable's cheapest good run and a thirteenth of Fable's best. Fable still edges it on detail and makes no visible errors, but it is close. And this is not a memorised pelican: it is general 3D-in-SVG ability from the cheapest tier of models.
MacBook

$0.07 · 9,544 tokens · 36s

$0.08 · 11,317 tokens · 47s

$0.16 · 21,051 tokens · 83s
Pelican

$0.10 · 12,996 tokens · 54s

$0.09 · 12,340 tokens · 50s

$0.09 · 11,520 tokens · 46s
Claude Fable 5 #
Anthropic · in Playcode since 2026-07-10
The best draughtsman here. Its xhigh MacBook won the launch sweep, and uncapped at max it produces the single best drawing in the benchmark - an isometric render with a notch, a dock and side ports, and no visible errors. You pay for it: $2.13 and nearly ten minutes for one laptop.
MacBook

$0.27 · 5,383 tokens · 68s

$0.73 · 14,554 tokens · 3.0 min

$2.13 · 42,560 tokens · 9.8 min
Pelican

$0.11 · 2,264 tokens · 26s

$0.12 · 2,400 tokens · 30s

$0.43 · 8,557 tokens · 112s
Grok 4.5 #
xAI · in Playcode since 2026-07-10
Barely tries. 668 output tokens on the MacBook - more logo than laptop - though that makes it the cheapest run in the benchmark at $0.004. Its pelican is fine. high is its ceiling.
MacBook

$0.004 · 668 tokens · 36s
Pelican

$0.009 · 1,460 tokens · 17s
GPT-5.6 Sol #
OpenAI · in Playcode since 2026-07-09
The most ambitious - more gradients, more perspective, more parts than anyone. At high that ambition bends the geometry; at xhigh it pulls together into a genuinely good laptop, admittedly more dark Lenovo than MacBook. Its party trick is the failure: it finishes the pelican at max, but on the MacBook it never returns at all.
MacBook

$0.26 · 8,551 tokens · 109s

$0.45 · 15,042 tokens · 3.5 min
Timed out. It did not return in 36 minutes, even with a 64K budget - exactly as in the first run.
- · - tokens · 36.6 min
Pelican

$0.09 · 3,088 tokens · 42s

$0.11 · 3,668 tokens · 51s

$0.44 · 14,690 tokens · 4.0 min
GPT-5.6 Terra #
OpenAI · in Playcode since 2026-07-09
Fast and reliable, and it does not know what a MacBook looks like. Even uncapped at max, with 29,834 tokens to spend, it draws a wide, thin generic ultrabook. The clearest evidence in this benchmark that budget does not buy understanding.
MacBook

$0.09 · 5,685 tokens · 34s

$0.17 · 11,650 tokens · 104s

$0.45 · 29,834 tokens · 4.6 min
Pelican

$0.05 · 3,083 tokens · 23s

$0.10 · 6,445 tokens · 54s
Truncated at the 16K cap.
$0.24 · 16,000 tokens · 2.5 min
Claude Sonnet 5 #
Anthropic · in Playcode since 2026-06-30
Broken geometry on the laptop, and the benchmark's worst runaway: at max it burned an entire 64K output budget without ever closing the SVG. Perfectly fine on a pelican, lost on an object you know well.
MacBook

$0.03 · 2,644 tokens · 29s

$0.12 · 12,011 tokens · 2.0 min
A genuine runaway: it burned the entire 64K budget and still never closed the SVG.
$0.66 · 65,538 tokens · 11.6 min
Pelican

$0.01 · 1,179 tokens · 10s

$0.02 · 1,769 tokens · 14s
Truncated: it hit the 16K output cap without ever closing the SVG.
$0.16 · 16,000 tokens · 3.0 min
Claude Opus 4.6 #
Anthropic · in Playcode since 2026-06-29
Rejects xhigh outright - the bug this benchmark found in its first ninety seconds. What it does draw is recognisable but loose, with parts in the wrong places under inspection.
MacBook

$0.20 · 8,156 tokens · 85s

$0.21 · 8,328 tokens · 82s

$0.19 · 7,767 tokens · 76s
Pelican

$0.05 · 2,091 tokens · 32s

$0.06 · 2,461 tokens · 33s

$0.05 · 1,942 tokens · 27s
Claude Opus 4.8 #
Anthropic · in Playcode since 2026-05-31
Cheap, fast, consistent - and consistently mid. Its three effort levels cost $0.05, $0.10 and $0.18 and produce what is recognisably the same drawing. The effort dial buys nothing here.
MacBook

$0.05 · 2,087 tokens · 21s

$0.10 · 3,912 tokens · 39s

$0.18 · 7,189 tokens · 65s
Pelican

$0.03 · 1,083 tokens · 14s

$0.03 · 1,183 tokens · 14s

$0.04 · 1,528 tokens · 16s
Claude Sonnet 4.6 #
Anthropic · in Playcode since 2026-03-02
Also rejects xhigh. Draws a laptop-shaped object with misplaced parts, and effort barely moves it: $0.11, $0.11, $0.10 across the ladder.
MacBook

$0.11 · 7,151 tokens · 82s

$0.11 · 7,340 tokens · 85s

$0.10 · 6,809 tokens · 77s
Pelican

$0.03 · 1,999 tokens · 25s

$0.03 · 1,854 tokens · 23s

$0.03 · 2,000 tokens · 24s
Gemini 3 Flash #
Google · in Playcode since 2026-01-20
The cheap-and-fast control, and a genuine surprise of the launch sweep: a simplified laptop with flat shading where nothing is broken, for $0.007. It seems to know what it can execute and stays inside that envelope.
MacBook

$0.007 · 2,348 tokens · 14s
Pelican

$0.007 · 2,328 tokens · 13s

$0.006 · 1,897 tokens · 12s

$0.008 · 2,536 tokens · 14s
What the effort dial buys #
The effort ladder was supposed to be the boring part. It produced the sharpest result of the benchmark, and it took two sweeps to state it correctly.
Our first sweep capped output at 16,000 tokens, and at max four models never finished. That looked like models falling off a cliff. So we re-ran the whole max column with a generous 64K budget - streaming, which Anthropic's SDK requires once a request could run past ten minutes. The honest picture is more interesting than the first one:
- Up to
xhigh, effort buys real quality on the strongest models. Fable's and Sol's best drawings both happen atxhigh, and the improvement overhighis visible detail. On mid models it buys nothing: Opus 4.8's three drawings cost $0.05, $0.10 and $0.18 and look identical. - Most of the truncations were our cap, not the model. Given 64K, Fable 5 finishes - and draws the best MacBook in the benchmark. Terra finishes too. Gemini 3.6 Flash finishes at max for $0.16.
- But finishing is not getting it right. Terra completes a full drawing at max and the form is still wrong - a wide, thin generic ultrabook. All the budget in the world does not buy an understanding of the object.
- Once in a while, max genuinely pays off. Kimi K3 is the clean case: its 30,579-token, 17.5-minute
maxMacBook is materially better than its ownhighone - correct 16-inch proportions, speaker grilles, no visible errors - and lands in Fable's quality tier for a fifth of the price. It is the exception that shows how rarely the top setting is worth its bill. - And some runs genuinely never end. Sonnet 5 - and now Opus 5 - burned an entire 64K budget without closing the SVG. Sol timed out after 36 minutes with nothing to show, twice. That is not a budget you can raise your way out of.
So max is mostly a trap. It reliably buys you a bigger bill and a longer wait; it occasionally buys a better drawing, and it sometimes buys nothing at all. If you run an effort dial in production, measure it on your own workload before you let users pay for the top setting.
What one drawing costs #
Because every cell is a real API call, we can price the same task across the market. This is every MacBook run, cheapest first - including the ones that billed and handed back no finished drawing.
| Model · effort | Cost per MacBook | vs cheapest |
|---|---|---|
| Grok 4.5 · high | $0.004 | 1.0x |
| Gemini 3 Flash · default | $0.007 | 1.7x |
| Claude Sonnet 5 · high | $0.03 | 6.2x |
| Claude Opus 4.8 · high | $0.05 | 12x |
| Gemini 3.6 Flash · low | $0.07 | 17x |
| Gemini 3.6 Flash · medium | $0.08 | 20x |
| GPT-5.6 Terra · high | $0.09 | 20x |
| Claude Opus 4.8 · xhigh | $0.10 | 23x |
| Claude Sonnet 4.6 · max | $0.10 | 24x |
| Claude Sonnet 4.6 · high | $0.11 | 25x |
| Claude Sonnet 4.6 · xhigh* | $0.11 | 26x |
| Claude Sonnet 5 · xhigh | $0.12 | 28x |
| Kimi K3 · high | $0.15 | 34x |
| Gemini 3.6 Flash · max (near-best, a fraction of the price) | $0.16 | 37x |
| GPT-5.6 Terra · xhigh | $0.17 | 41x |
| Claude Opus 4.8 · max | $0.18 | 42x |
| Claude Opus 4.6 · max | $0.19 | 45x |
| Claude Opus 4.6 · high | $0.20 | 47x |
| Claude Opus 4.6 · xhigh* | $0.21 | 49x |
| GPT-5.6 Sol · high | $0.26 | 60x |
| Claude Fable 5 · high | $0.27 | 63x |
| GPT-5.6 Terra · max | $0.45 | 104x |
| GPT-5.6 Sol · xhigh | $0.45 | 105x |
| Kimi K3 · max | $0.46 | 107x |
| Claude Opus 5 · high NEW | $0.50 | 117x |
| Claude Sonnet 5 · max | $0.66· no drawing | 152x |
| Claude Fable 5 · xhigh | $0.73 | 169x |
| Claude Opus 5 · xhigh NEW | $1.18 | 275x |
| Claude Opus 5 · max NEW | $1.60· no drawing | 372x |
| Claude Fable 5 · max (the best drawing) | $2.13 | 495x |
| GPT-5.6 Sol · max | not billed· no drawing | - |
The best drawing costs about 495x the cheapest passable one, and the cheapest entries are not embarrassing. Note the greyed rows: you can pay full price and receive nothing. They billed $2.26 for files that were never closed - and Sol's max is not even counted there, because it never returned to bill. Whether Fable's detail is worth 13x Gemini 3.6's coherence depends entirely on what you are building - but you cannot even ask that question from a $/Mtok pricing page. (Why the sticker price misleads across vendors is its own story: The Same TypeScript Costs 73% More on Claude Than on GPT.)
The bug this found #
The first full run crashed two cells with an API error we had never seen in production: This model does not support effort level 'xhigh'. Anthropic's xhigh exists only on the newest generation (Sonnet 5, Opus 4.8, Fable 5); Opus 4.6 and Sonnet 4.6 reject it, and their supported list is not even a superset-consistent scale across generations. Our product maps a user-facing quality dial onto these levels, which means one combination of settings produced a hard error on every message. The benchmark surfaced it in ninety seconds and we shipped the fix the same night. A static leaderboard would never have caught it.
The second sweep found a second one, in our own harness: Gemini's effort dial had been silently doing nothing. Gemini exposes thinking_level rather than a reasoning-effort parameter, so every Gemini run in the launch sweep was quietly at the default. Wiring it up is why Gemini 3 Flash finally has an effort ladder in the matrix above - and why Gemini 3.6 Flash could be measured properly at all.
Try it yourself #
The method is deliberately reproducible. The prompt, ready to paste into any model:
Create an SVG 3D model of a MacBook Pro 16 in a dark color. Respond with ONLY the SVG markup, no explanation.Or swap in any object your audience knows intimately - a Coke can, a Vespa, a Stratocaster. The only requirements are that the ground truth lives in your reader's head and that style cannot hide broken geometry. We re-run this gallery when new frontier models ship; the pelican had a good run, and the MacBook has years of discrimination left in it.
Method notes: runs on 2026-07-13/14, 2026-07-21/22 and 2026-07-24 via each provider's public API, sampling left at provider defaults, reasoning effort set through each provider's native parameter (Anthropic output_config.effort, OpenAI reasoning.effort, xAI reasoning_effort, Google thinking_level, Moonshot reasoning_effort). Output capped at 16,000 tokens in the first sweep and 64K thereafter; the cap is stated wherever it changed a result. Exact model ids: claude-fable-5, claude-opus-5, claude-opus-4-8, claude-opus-4-6, claude-sonnet-5, claude-sonnet-4-6, gpt-5.6-sol, gpt-5.6-terra, grok-4.5, gemini-3-flash-preview, gemini-3.6-flash and kimi-k3. Costs come from the providers' own usage-reported token counts at list prices; total measured spend across every run shown here is $13.80. Drawings are unedited; failures are reported as failures rather than re-rolled. The measurements are ours; the prose was drafted with AI assistance and edited by a human.
Playcode keeps all of these models one click apart, with the same effort dial we benchmarked here - so you can run your own MacBook test on a real project instead of taking our word for it. Try it at playcode.io.