MacBook SVG Benchmark: 12 AI Models | Playcode Blog

24 min read Original article ↗

TL;DR

  • Every frontier model passes the pelican-on-a-bicycle test now. It is in every training corpus and was never precisely gradable, so it stopped separating anyone.
  • Our replacement: draw a MacBook Pro 16 in SVG, one shot. Everyone knows the exact proportions, so a 5% error is visible instantly.
  • Claude Fable 5 draws the best one. Gemini 3.6 Flash gets startlingly close for $0.16 - a thirteenth of Fable's best run.
  • New: Claude Opus 5 (Anthropic) draws the most detailed Anthropic MacBook here - but the geometry is broken at every effort level, it only looks the part, and thinking-on-by-default burns 20-47K tokens per drawing (10-25x Opus 4.8's cost); its max run spends the whole budget thinking and never draws the laptop.
  • The max effort setting is mostly a trap: it buys cost and latency, sometimes a run that never terminates, and never an understanding of the object.
  • 64 runs across 12 models, both tasks, every effort level. Unedited output, providers' own billing.

NEW2026-07-24: Added Claude Opus 5 (Anthropic) at high/xhigh/max, both tasks - the most detailed Anthropic laptop here, but the geometry is broken at every level (skewed at high, broken isometry at xhigh, no laptop at all at max), and thinking-on-by-default makes it 10-25x pricier per drawing than Opus 4.8.

For two years the best quick test of a new model was Simon Willison's pelican: ask for a pelican riding a bicycle as SVG and look at the result. It worked because it could not be faked. It no longer separates anyone - so we built the test we actually wanted: draw a MacBook Pro 16 in SVG, in a dark colour, one shot. Everyone reading this knows exactly what a MacBook looks like, so everyone reading this can grade it. Below is every run we have: 12 models, both tasks, the full effort ladder, 64 runs, unedited.

Grid of SVG MacBook Pro drawings produced by frontier AI models, ranging from a correctly assembled dark laptop to broken abstract shapes

Why the pelican had to go #

A benchmark is useful exactly as long as it separates models. We ran the pelican as a control through the identical harness - same models, same effort ladder, same one-shot rule - and every single model produced a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading. Two problems got it there:

  • Everyone has trained on it. It is the most famous drawing prompt in AI. Three years of pelicans and public commentary about them sit in every training corpus. A good pelican no longer tells you whether the model understands geometry or has simply seen ten thousand graded pelicans.
  • It never had a gradation. Nobody knows how long a pelican's beak should be relative to its wingspan, so grading tops out at "recognisable". There is no scale on which one passing pelican beats another.

A MacBook Pro fixes both. Hundreds of millions of people have stared at one for thousands of hours, so the lid, the keyboard grid, the notch, the hinge and the trackpad have a precise shared ground truth - your eye flags a small error instantly. And it has a real difficulty ladder: dozens of parts plus a 3D perspective view. Models can now fail at many different heights instead of clustering at "recognisable". The numbers say the same thing: the models spend roughly 4x fewer output tokens on a pelican than on the MacBook (Opus 4.6: 2,091 vs 8,156), and cranking effort barely moves the pelican at all.

How we ran it #

Create an SVG 3D model of a MacBook Pro 16 in a dark color. Respond with ONLY the SVG markup, no explanation.

One API call per cell. No retries, no system prompt, no examples, no cherry-picking - the first answer is the answer. Each model runs the effort ladder its provider supports (high, xhigh, max); Grok 4.5 tops out at high, Gemini 3.6 Flash exposes low/medium instead of xhigh, and Kimi K3 runs low/high/max with no xhigh. Costs are the providers' own usage-reported billing. Runs marked * were requested as xhigh and clamped to high by the model's generation.

On the images. Each drawing is a browser render of the model's unedited SVG - nothing touched up, and the raw SVGs stay in the repo. We rasterise deliberately: the models all reach for the same generic ids (#screenGrad, #screenClip, #shadow), so inlining fifty of them into one page makes the browser resolve one model's clip-path into another model's drawing. Rendering each SVG in its own document is the only way to show you what the model actually drew.

On reproducibility. We re-ran the pelican ladder nine days later on a different output budget; token counts landed within about 10% of the originals (Fable at high: 2,264 then 2,002). Each run below records the sweep it came from, so mixed dates stay visible rather than quietly blended.

Every run #

The whole matrix, both tasks. Filter to one task or vendor, sort by cost to line them up cheapest-to-priciest, or show only the newest model.

64 of 64 runs

NEW

Claude Opus 5 macbook at high effort

Claude Opus 5 · macbook · high
$0.50 · 20,145 tokens · 4.3 min
Anthropic · run 2026-07-24

All the detail is here - notch, dock with app icons, a code window, "MacBook Pro" under the screen - but the base is skewed: the deck rakes and runs too deep, in a slightly different perspective from the lid. $0.50 and 20K output tokens, because thinking is on by default.

NEW

Claude Opus 5 macbook at xhigh effort

Claude Opus 5 · macbook · xhigh
$1.18 · 47,237 tokens · 8.5 min
Anthropic · run 2026-07-24

More detail and better materials than high, but the isometry is broken: the screen and the keyboard deck sit in different perspectives, so it looks slick at a glance and falls apart on a second look. 47K output tokens and $1.18 for it.

NEW

Runaway. Thinking is on by default, so it spent almost the entire 64K budget reasoning and emitted only the background glow before the cap - no laptop at all, $1.60 billed for nothing. The same failure mode as Sonnet 5.

Claude Opus 5 · macbook · max
$1.60 · 64,000 tokens · 13.3 min
Anthropic · run 2026-07-24
NEW

Claude Opus 5 pelican at high effort

Claude Opus 5 · pelican · high
$0.06 · 2,290 tokens · 28s
Anthropic · run 2026-07-24

Clean and correct first try - white body, big orange pouch, both wheels properly spoked, legs on the pedals.

NEW

Claude Opus 5 pelican at xhigh effort

Claude Opus 5 · pelican · xhigh
$0.09 · 3,681 tokens · 46s
Anthropic · run 2026-07-24
NEW

Claude Opus 5 pelican at max effort

Claude Opus 5 · pelican · max
$0.15 · 5,957 tokens · 76s
Anthropic · run 2026-07-24

The dial that does nothing on other pelicans does something here: max adds a dynamic pose, speed lines and a curved handlebar over high. On a saturated task, that is polish, not correctness.

Kimi K3 macbook at high effort

Kimi K3 · macbook · high
$0.15 · 9,846 tokens · 5.0 min
Moonshot AI · run 2026-07-24

A proper dark MacBook - dock, notch, keyboard, trackpad - already here for $0.15, a hair over Gemini 3.6 Flash at max.

Kimi K3 macbook at max effort

Kimi K3 · macbook · max
$0.46 · 30,579 tokens · 17.5 min
Moonshot AI · run 2026-07-24

Max earns it here - one of the best laptops in the benchmark: correct 16-inch proportions, speaker grilles, a code window, no visible errors. Fable-tier for a fifth of the price, but 30,579 tokens and 17.5 minutes.

Kimi K3 pelican at high effort

Kimi K3 · pelican · high
$0.05 · 3,425 tokens · 108s
Moonshot AI · run 2026-07-24

Aces it first try - clean bike, big orange pouch, a crest and motion lines.

Kimi K3 pelican at max effort

Kimi K3 · pelican · max
$0.09 · 6,183 tokens · 3.5 min
Moonshot AI · run 2026-07-24

6,183 tokens against high's 3,425 for a pelican that is no better - the effort dial doing nothing again.

Gemini 3.6 Flash macbook at low effort

Gemini 3.6 Flash · macbook · low
$0.07 · 9,544 tokens · 36s
Google · run 2026-07-21

Gemini 3.6 Flash macbook at medium effort

Gemini 3.6 Flash · macbook · medium
$0.08 · 11,317 tokens · 47s
Google · run 2026-07-21

Gemini 3.6 Flash macbook at max effort

Gemini 3.6 Flash · macbook · max
$0.16 · 21,051 tokens · 83s
Google · run 2026-07-21

Max effort, complete: a detailed dark laptop in a proper three-quarter view, for $0.16.

Gemini 3.6 Flash pelican at high effort

Gemini 3.6 Flash · pelican · high
$0.10 · 12,996 tokens · 54s
Google · run 2026-07-21

Gemini 3.6 Flash pelican at xhigh effort

Gemini 3.6 Flash · pelican · xhigh
$0.09 · 12,340 tokens · 50s
Google · run 2026-07-21

Gemini 3.6 Flash pelican at max effort

Gemini 3.6 Flash · pelican · max
$0.09 · 11,520 tokens · 46s
Google · run 2026-07-21

Claude Fable 5 macbook at high effort

Claude Fable 5 · macbook · high
$0.27 · 5,383 tokens · 68s
Anthropic · run 2026-07-13

Claude Fable 5 macbook at xhigh effort

Claude Fable 5 · macbook · xhigh
$0.73 · 14,554 tokens · 3.0 min
Anthropic · run 2026-07-13

Winner of the launch sweep - the most detailed MacBook anyone drew inside the 16K cap.

Claude Fable 5 macbook at max effort

Claude Fable 5 · macbook · max
$2.13 · 42,560 tokens · 9.8 min
Anthropic · run 2026-07-21

Uncapped (64K, streamed): the best drawing in the benchmark - isometric, dock, side ports - for $2.13 and 9.8 minutes.

Grok 4.5 macbook at high effort

Grok 4.5 · macbook · high
$0.004 · 668 tokens · 36s
xAI · run 2026-07-13

high is its ceiling. 668 output tokens - it barely tried.

Claude Fable 5 pelican at high effort

Claude Fable 5 · pelican · high
$0.11 · 2,264 tokens · 26s
Anthropic · run 2026-07-13

Claude Fable 5 pelican at xhigh effort

Claude Fable 5 · pelican · xhigh
$0.12 · 2,400 tokens · 30s
Anthropic · run 2026-07-13

Claude Fable 5 pelican at max effort

Claude Fable 5 · pelican · max
$0.43 · 8,557 tokens · 112s
Anthropic · run 2026-07-13

Grok 4.5 pelican at high effort

Grok 4.5 · pelican · high
$0.009 · 1,460 tokens · 17s
xAI · run 2026-07-13

GPT-5.6 Sol macbook at high effort

GPT-5.6 Sol · macbook · high
$0.26 · 8,551 tokens · 109s
OpenAI · run 2026-07-13

GPT-5.6 Sol macbook at xhigh effort

GPT-5.6 Sol · macbook · xhigh
$0.45 · 15,042 tokens · 3.5 min
OpenAI · run 2026-07-13

Timed out. It did not return in 36 minutes, even with a 64K budget - exactly as in the first run.

GPT-5.6 Sol · macbook · max
- · - tokens · 36.6 min
OpenAI · run 2026-07-21

GPT-5.6 Terra macbook at high effort

GPT-5.6 Terra · macbook · high
$0.09 · 5,685 tokens · 34s
OpenAI · run 2026-07-13

GPT-5.6 Terra macbook at xhigh effort

GPT-5.6 Terra · macbook · xhigh
$0.17 · 11,650 tokens · 104s
OpenAI · run 2026-07-13

GPT-5.6 Terra macbook at max effort

GPT-5.6 Terra · macbook · max
$0.45 · 29,834 tokens · 4.6 min
OpenAI · run 2026-07-21

Uncapped it finishes - and the form is still wrong. Budget does not buy understanding.

GPT-5.6 Sol pelican at high effort

GPT-5.6 Sol · pelican · high
$0.09 · 3,088 tokens · 42s
OpenAI · run 2026-07-13

GPT-5.6 Sol pelican at xhigh effort

GPT-5.6 Sol · pelican · xhigh
$0.11 · 3,668 tokens · 51s
OpenAI · run 2026-07-13

GPT-5.6 Sol pelican at max effort

GPT-5.6 Sol · pelican · max
$0.44 · 14,690 tokens · 4.0 min
OpenAI · run 2026-07-13

Sol does finish at max here. The task it cannot finish is the MacBook.

GPT-5.6 Terra pelican at high effort

GPT-5.6 Terra · pelican · high
$0.05 · 3,083 tokens · 23s
OpenAI · run 2026-07-13

GPT-5.6 Terra pelican at xhigh effort

GPT-5.6 Terra · pelican · xhigh
$0.10 · 6,445 tokens · 54s
OpenAI · run 2026-07-13

Truncated at the 16K cap.

GPT-5.6 Terra · pelican · max
$0.24 · 16,000 tokens · 2.5 min
OpenAI · run 2026-07-13

Claude Sonnet 5 macbook at high effort

Claude Sonnet 5 · macbook · high
$0.03 · 2,644 tokens · 29s
Anthropic · run 2026-07-13

Claude Sonnet 5 macbook at xhigh effort

Claude Sonnet 5 · macbook · xhigh
$0.12 · 12,011 tokens · 2.0 min
Anthropic · run 2026-07-13

A genuine runaway: it burned the entire 64K budget and still never closed the SVG.

Claude Sonnet 5 · macbook · max
$0.66 · 65,538 tokens · 11.6 min
Anthropic · run 2026-07-21

Claude Sonnet 5 pelican at high effort

Claude Sonnet 5 · pelican · high
$0.01 · 1,179 tokens · 10s
Anthropic · run 2026-07-13

Claude Sonnet 5 pelican at xhigh effort

Claude Sonnet 5 · pelican · xhigh
$0.02 · 1,769 tokens · 14s
Anthropic · run 2026-07-13

Truncated: it hit the 16K output cap without ever closing the SVG.

Claude Sonnet 5 · pelican · max
$0.16 · 16,000 tokens · 3.0 min
Anthropic · run 2026-07-13

Claude Opus 4.6 macbook at high effort

Claude Opus 4.6 · macbook · high
$0.20 · 8,156 tokens · 85s
Anthropic · run 2026-07-13

Claude Opus 4.6 macbook at xhigh effort

Claude Opus 4.6 · macbook · xhigh*
$0.21 · 8,328 tokens · 82s
Anthropic · run 2026-07-13

Claude Opus 4.6 macbook at max effort

Claude Opus 4.6 · macbook · max
$0.19 · 7,767 tokens · 76s
Anthropic · run 2026-07-13

Claude Opus 4.6 pelican at high effort

Claude Opus 4.6 · pelican · high
$0.05 · 2,091 tokens · 32s
Anthropic · run 2026-07-13

Claude Opus 4.6 pelican at xhigh effort

Claude Opus 4.6 · pelican · xhigh*
$0.06 · 2,461 tokens · 33s
Anthropic · run 2026-07-13

Claude Opus 4.6 pelican at max effort

Claude Opus 4.6 · pelican · max
$0.05 · 1,942 tokens · 27s
Anthropic · run 2026-07-13

Claude Opus 4.8 macbook at high effort

Claude Opus 4.8 · macbook · high
$0.05 · 2,087 tokens · 21s
Anthropic · run 2026-07-13

Claude Opus 4.8 macbook at xhigh effort

Claude Opus 4.8 · macbook · xhigh
$0.10 · 3,912 tokens · 39s
Anthropic · run 2026-07-13

Claude Opus 4.8 macbook at max effort

Claude Opus 4.8 · macbook · max
$0.18 · 7,189 tokens · 65s
Anthropic · run 2026-07-13

Claude Opus 4.8 pelican at high effort

Claude Opus 4.8 · pelican · high
$0.03 · 1,083 tokens · 14s
Anthropic · run 2026-07-13

Claude Opus 4.8 pelican at xhigh effort

Claude Opus 4.8 · pelican · xhigh
$0.03 · 1,183 tokens · 14s
Anthropic · run 2026-07-13

Claude Opus 4.8 pelican at max effort

Claude Opus 4.8 · pelican · max
$0.04 · 1,528 tokens · 16s
Anthropic · run 2026-07-13

Claude Sonnet 4.6 macbook at high effort

Claude Sonnet 4.6 · macbook · high
$0.11 · 7,151 tokens · 82s
Anthropic · run 2026-07-13

Claude Sonnet 4.6 macbook at xhigh effort

Claude Sonnet 4.6 · macbook · xhigh*
$0.11 · 7,340 tokens · 85s
Anthropic · run 2026-07-13

Claude Sonnet 4.6 macbook at max effort

Claude Sonnet 4.6 · macbook · max
$0.10 · 6,809 tokens · 77s
Anthropic · run 2026-07-13

Claude Sonnet 4.6 pelican at high effort

Claude Sonnet 4.6 · pelican · high
$0.03 · 1,999 tokens · 25s
Anthropic · run 2026-07-13

Claude Sonnet 4.6 pelican at xhigh effort

Claude Sonnet 4.6 · pelican · xhigh*
$0.03 · 1,854 tokens · 23s
Anthropic · run 2026-07-13

Claude Sonnet 4.6 pelican at max effort

Claude Sonnet 4.6 · pelican · max
$0.03 · 2,000 tokens · 24s
Anthropic · run 2026-07-13

Gemini 3 Flash macbook at default effort

Gemini 3 Flash · macbook · default
$0.007 · 2,348 tokens · 14s
Google · run 2026-07-13

Gemini 3 Flash pelican at default effort

Gemini 3 Flash · pelican · default
$0.007 · 2,328 tokens · 13s
Google · run 2026-07-13

Gemini 3 Flash pelican at high effort

Gemini 3 Flash · pelican · high
$0.006 · 1,897 tokens · 12s
Google · run 2026-07-21

Only possible after we wired Gemini to thinking_level - this model had no working effort dial in the launch sweep.

Gemini 3 Flash pelican at xhigh effort

Gemini 3 Flash · pelican · xhigh
$0.008 · 2,536 tokens · 14s
Google · run 2026-07-21

Model by model #

Each model with both of its drawings side by side - the pelican it aces and the MacBook that actually grades it. Newest first.

Claude Opus 5 NEW#

Anthropic · in Playcode since 2026-07-24

The most detailed, ambitious Anthropic laptop here - notch, dock, traffic lights, a code window, convincing materials - and the first Anthropic model whose effort dial actually changes the drawing instead of redrawing the same thing. But it fails the one thing the MacBook test is for: the geometry. At high the base is skewed; at xhigh the isometry is broken outright - the screen and the keyboard deck sit in different perspectives, so it reads slick at a glance and falls apart on a second look. And max is the worst of the three: thinking is on by default, so it spent almost the entire 64K budget reasoning and emitted only the background glow before the cap - no laptop at all, $1.60 for nothing. A clear step up from Opus 4.8's flat mid laptop in detail and effort-response - but what separates Fable is that it makes no visual errors, where Opus 5 only looks the part.

MacBook

Claude Opus 5 macbook at high effort

high
$0.50 · 20,145 tokens · 4.3 min

Claude Opus 5 macbook at xhigh effort

xhigh
$1.18 · 47,237 tokens · 8.5 min

Runaway. Thinking is on by default, so it spent almost the entire 64K budget reasoning and emitted only the background glow before the cap - no laptop at all, $1.60 billed for nothing. The same failure mode as Sonnet 5.

max
$1.60 · 64,000 tokens · 13.3 min

Pelican

Claude Opus 5 pelican at high effort

high
$0.06 · 2,290 tokens · 28s

Claude Opus 5 pelican at xhigh effort

xhigh
$0.09 · 3,681 tokens · 46s

Claude Opus 5 pelican at max effort

max
$0.15 · 5,957 tokens · 76s

Kimi K3 #

Moonshot AI · in Playcode since 2026-07-23

The new challenger, and a serious one. It draws a proper dark MacBook - dock, notch, speaker grilles - at both levels, and at max it produces one of the best laptops in the benchmark: a clean render with correct 16-inch proportions and no visible errors, for $0.46. That is a fifth of Fable's price for a drawing in the same class. The catch is the clock - that run took 17.5 minutes - and, as ever, the pelican it already aces at high gains nothing from max.

MacBook

Kimi K3 macbook at high effort

high
$0.15 · 9,846 tokens · 5.0 min

Kimi K3 macbook at max effort

max
$0.46 · 30,579 tokens · 17.5 min

Pelican

Kimi K3 pelican at high effort

high
$0.05 · 3,425 tokens · 108s

Kimi K3 pelican at max effort

max
$0.09 · 6,183 tokens · 3.5 min

Gemini 3.6 Flash #

Google · in Playcode since 2026-07-21

The surprise of the whole benchmark. Google's newest Flash is the first Gemini here with a working effort dial, and at max it draws a near-perfect 3D MacBook for $0.16 - a fifth of Fable's cheapest good run and a thirteenth of Fable's best. Fable still edges it on detail and makes no visible errors, but it is close. And this is not a memorised pelican: it is general 3D-in-SVG ability from the cheapest tier of models.

MacBook

Gemini 3.6 Flash macbook at low effort

low
$0.07 · 9,544 tokens · 36s

Gemini 3.6 Flash macbook at medium effort

medium
$0.08 · 11,317 tokens · 47s

Gemini 3.6 Flash macbook at max effort

max
$0.16 · 21,051 tokens · 83s

Pelican

Gemini 3.6 Flash pelican at high effort

high
$0.10 · 12,996 tokens · 54s

Gemini 3.6 Flash pelican at xhigh effort

xhigh
$0.09 · 12,340 tokens · 50s

Gemini 3.6 Flash pelican at max effort

max
$0.09 · 11,520 tokens · 46s

Claude Fable 5 #

Anthropic · in Playcode since 2026-07-10

The best draughtsman here. Its xhigh MacBook won the launch sweep, and uncapped at max it produces the single best drawing in the benchmark - an isometric render with a notch, a dock and side ports, and no visible errors. You pay for it: $2.13 and nearly ten minutes for one laptop.

MacBook

Claude Fable 5 macbook at high effort

high
$0.27 · 5,383 tokens · 68s

Claude Fable 5 macbook at xhigh effort

xhigh
$0.73 · 14,554 tokens · 3.0 min

Claude Fable 5 macbook at max effort

max
$2.13 · 42,560 tokens · 9.8 min

Pelican

Claude Fable 5 pelican at high effort

high
$0.11 · 2,264 tokens · 26s

Claude Fable 5 pelican at xhigh effort

xhigh
$0.12 · 2,400 tokens · 30s

Claude Fable 5 pelican at max effort

max
$0.43 · 8,557 tokens · 112s

Grok 4.5 #

xAI · in Playcode since 2026-07-10

Barely tries. 668 output tokens on the MacBook - more logo than laptop - though that makes it the cheapest run in the benchmark at $0.004. Its pelican is fine. high is its ceiling.

MacBook

Grok 4.5 macbook at high effort

high
$0.004 · 668 tokens · 36s

Pelican

Grok 4.5 pelican at high effort

high
$0.009 · 1,460 tokens · 17s

GPT-5.6 Sol #

OpenAI · in Playcode since 2026-07-09

The most ambitious - more gradients, more perspective, more parts than anyone. At high that ambition bends the geometry; at xhigh it pulls together into a genuinely good laptop, admittedly more dark Lenovo than MacBook. Its party trick is the failure: it finishes the pelican at max, but on the MacBook it never returns at all.

MacBook

GPT-5.6 Sol macbook at high effort

high
$0.26 · 8,551 tokens · 109s

GPT-5.6 Sol macbook at xhigh effort

xhigh
$0.45 · 15,042 tokens · 3.5 min

Timed out. It did not return in 36 minutes, even with a 64K budget - exactly as in the first run.

max
- · - tokens · 36.6 min

Pelican

GPT-5.6 Sol pelican at high effort

high
$0.09 · 3,088 tokens · 42s

GPT-5.6 Sol pelican at xhigh effort

xhigh
$0.11 · 3,668 tokens · 51s

GPT-5.6 Sol pelican at max effort

max
$0.44 · 14,690 tokens · 4.0 min

GPT-5.6 Terra #

OpenAI · in Playcode since 2026-07-09

Fast and reliable, and it does not know what a MacBook looks like. Even uncapped at max, with 29,834 tokens to spend, it draws a wide, thin generic ultrabook. The clearest evidence in this benchmark that budget does not buy understanding.

MacBook

GPT-5.6 Terra macbook at high effort

high
$0.09 · 5,685 tokens · 34s

GPT-5.6 Terra macbook at xhigh effort

xhigh
$0.17 · 11,650 tokens · 104s

GPT-5.6 Terra macbook at max effort

max
$0.45 · 29,834 tokens · 4.6 min

Pelican

GPT-5.6 Terra pelican at high effort

high
$0.05 · 3,083 tokens · 23s

GPT-5.6 Terra pelican at xhigh effort

xhigh
$0.10 · 6,445 tokens · 54s

Truncated at the 16K cap.

max
$0.24 · 16,000 tokens · 2.5 min

Claude Sonnet 5 #

Anthropic · in Playcode since 2026-06-30

Broken geometry on the laptop, and the benchmark's worst runaway: at max it burned an entire 64K output budget without ever closing the SVG. Perfectly fine on a pelican, lost on an object you know well.

MacBook

Claude Sonnet 5 macbook at high effort

high
$0.03 · 2,644 tokens · 29s

Claude Sonnet 5 macbook at xhigh effort

xhigh
$0.12 · 12,011 tokens · 2.0 min

A genuine runaway: it burned the entire 64K budget and still never closed the SVG.

max
$0.66 · 65,538 tokens · 11.6 min

Pelican

Claude Sonnet 5 pelican at high effort

high
$0.01 · 1,179 tokens · 10s

Claude Sonnet 5 pelican at xhigh effort

xhigh
$0.02 · 1,769 tokens · 14s

Truncated: it hit the 16K output cap without ever closing the SVG.

max
$0.16 · 16,000 tokens · 3.0 min

Claude Opus 4.6 #

Anthropic · in Playcode since 2026-06-29

Rejects xhigh outright - the bug this benchmark found in its first ninety seconds. What it does draw is recognisable but loose, with parts in the wrong places under inspection.

MacBook

Claude Opus 4.6 macbook at high effort

high
$0.20 · 8,156 tokens · 85s

Claude Opus 4.6 macbook at xhigh effort

xhigh*
$0.21 · 8,328 tokens · 82s

Claude Opus 4.6 macbook at max effort

max
$0.19 · 7,767 tokens · 76s

Pelican

Claude Opus 4.6 pelican at high effort

high
$0.05 · 2,091 tokens · 32s

Claude Opus 4.6 pelican at xhigh effort

xhigh*
$0.06 · 2,461 tokens · 33s

Claude Opus 4.6 pelican at max effort

max
$0.05 · 1,942 tokens · 27s

Claude Opus 4.8 #

Anthropic · in Playcode since 2026-05-31

Cheap, fast, consistent - and consistently mid. Its three effort levels cost $0.05, $0.10 and $0.18 and produce what is recognisably the same drawing. The effort dial buys nothing here.

MacBook

Claude Opus 4.8 macbook at high effort

high
$0.05 · 2,087 tokens · 21s

Claude Opus 4.8 macbook at xhigh effort

xhigh
$0.10 · 3,912 tokens · 39s

Claude Opus 4.8 macbook at max effort

max
$0.18 · 7,189 tokens · 65s

Pelican

Claude Opus 4.8 pelican at high effort

high
$0.03 · 1,083 tokens · 14s

Claude Opus 4.8 pelican at xhigh effort

xhigh
$0.03 · 1,183 tokens · 14s

Claude Opus 4.8 pelican at max effort

max
$0.04 · 1,528 tokens · 16s

Claude Sonnet 4.6 #

Anthropic · in Playcode since 2026-03-02

Also rejects xhigh. Draws a laptop-shaped object with misplaced parts, and effort barely moves it: $0.11, $0.11, $0.10 across the ladder.

MacBook

Claude Sonnet 4.6 macbook at high effort

high
$0.11 · 7,151 tokens · 82s

Claude Sonnet 4.6 macbook at xhigh effort

xhigh*
$0.11 · 7,340 tokens · 85s

Claude Sonnet 4.6 macbook at max effort

max
$0.10 · 6,809 tokens · 77s

Pelican

Claude Sonnet 4.6 pelican at high effort

high
$0.03 · 1,999 tokens · 25s

Claude Sonnet 4.6 pelican at xhigh effort

xhigh*
$0.03 · 1,854 tokens · 23s

Claude Sonnet 4.6 pelican at max effort

max
$0.03 · 2,000 tokens · 24s

Gemini 3 Flash #

Google · in Playcode since 2026-01-20

The cheap-and-fast control, and a genuine surprise of the launch sweep: a simplified laptop with flat shading where nothing is broken, for $0.007. It seems to know what it can execute and stays inside that envelope.

MacBook

Gemini 3 Flash macbook at default effort

default
$0.007 · 2,348 tokens · 14s

Pelican

Gemini 3 Flash pelican at default effort

default
$0.007 · 2,328 tokens · 13s

Gemini 3 Flash pelican at high effort

high
$0.006 · 1,897 tokens · 12s

Gemini 3 Flash pelican at xhigh effort

xhigh
$0.008 · 2,536 tokens · 14s

What the effort dial buys #

The effort ladder was supposed to be the boring part. It produced the sharpest result of the benchmark, and it took two sweeps to state it correctly.

Our first sweep capped output at 16,000 tokens, and at max four models never finished. That looked like models falling off a cliff. So we re-ran the whole max column with a generous 64K budget - streaming, which Anthropic's SDK requires once a request could run past ten minutes. The honest picture is more interesting than the first one:

  • Up to xhigh, effort buys real quality on the strongest models. Fable's and Sol's best drawings both happen at xhigh, and the improvement over high is visible detail. On mid models it buys nothing: Opus 4.8's three drawings cost $0.05, $0.10 and $0.18 and look identical.
  • Most of the truncations were our cap, not the model. Given 64K, Fable 5 finishes - and draws the best MacBook in the benchmark. Terra finishes too. Gemini 3.6 Flash finishes at max for $0.16.
  • But finishing is not getting it right. Terra completes a full drawing at max and the form is still wrong - a wide, thin generic ultrabook. All the budget in the world does not buy an understanding of the object.
  • Once in a while, max genuinely pays off. Kimi K3 is the clean case: its 30,579-token, 17.5-minute max MacBook is materially better than its own high one - correct 16-inch proportions, speaker grilles, no visible errors - and lands in Fable's quality tier for a fifth of the price. It is the exception that shows how rarely the top setting is worth its bill.
  • And some runs genuinely never end. Sonnet 5 - and now Opus 5 - burned an entire 64K budget without closing the SVG. Sol timed out after 36 minutes with nothing to show, twice. That is not a budget you can raise your way out of.

So max is mostly a trap. It reliably buys you a bigger bill and a longer wait; it occasionally buys a better drawing, and it sometimes buys nothing at all. If you run an effort dial in production, measure it on your own workload before you let users pay for the top setting.

What one drawing costs #

Because every cell is a real API call, we can price the same task across the market. This is every MacBook run, cheapest first - including the ones that billed and handed back no finished drawing.

Model · effortCost per MacBookvs cheapest
Grok 4.5 · high $0.0041.0x
Gemini 3 Flash · default $0.0071.7x
Claude Sonnet 5 · high $0.036.2x
Claude Opus 4.8 · high $0.0512x
Gemini 3.6 Flash · low $0.0717x
Gemini 3.6 Flash · medium $0.0820x
GPT-5.6 Terra · high $0.0920x
Claude Opus 4.8 · xhigh $0.1023x
Claude Sonnet 4.6 · max $0.1024x
Claude Sonnet 4.6 · high $0.1125x
Claude Sonnet 4.6 · xhigh* $0.1126x
Claude Sonnet 5 · xhigh $0.1228x
Kimi K3 · high $0.1534x
Gemini 3.6 Flash · max (near-best, a fraction of the price)$0.1637x
GPT-5.6 Terra · xhigh $0.1741x
Claude Opus 4.8 · max $0.1842x
Claude Opus 4.6 · max $0.1945x
Claude Opus 4.6 · high $0.2047x
Claude Opus 4.6 · xhigh* $0.2149x
GPT-5.6 Sol · high $0.2660x
Claude Fable 5 · high $0.2763x
GPT-5.6 Terra · max $0.45104x
GPT-5.6 Sol · xhigh $0.45105x
Kimi K3 · max $0.46107x
Claude Opus 5 · high NEW$0.50117x
Claude Sonnet 5 · max $0.66· no drawing152x
Claude Fable 5 · xhigh $0.73169x
Claude Opus 5 · xhigh NEW$1.18275x
Claude Opus 5 · max NEW$1.60· no drawing372x
Claude Fable 5 · max (the best drawing)$2.13495x
GPT-5.6 Sol · max not billed· no drawing-

The best drawing costs about 495x the cheapest passable one, and the cheapest entries are not embarrassing. Note the greyed rows: you can pay full price and receive nothing. They billed $2.26 for files that were never closed - and Sol's max is not even counted there, because it never returned to bill. Whether Fable's detail is worth 13x Gemini 3.6's coherence depends entirely on what you are building - but you cannot even ask that question from a $/Mtok pricing page. (Why the sticker price misleads across vendors is its own story: The Same TypeScript Costs 73% More on Claude Than on GPT.)

The bug this found #

The first full run crashed two cells with an API error we had never seen in production: This model does not support effort level 'xhigh'. Anthropic's xhigh exists only on the newest generation (Sonnet 5, Opus 4.8, Fable 5); Opus 4.6 and Sonnet 4.6 reject it, and their supported list is not even a superset-consistent scale across generations. Our product maps a user-facing quality dial onto these levels, which means one combination of settings produced a hard error on every message. The benchmark surfaced it in ninety seconds and we shipped the fix the same night. A static leaderboard would never have caught it.

The second sweep found a second one, in our own harness: Gemini's effort dial had been silently doing nothing. Gemini exposes thinking_level rather than a reasoning-effort parameter, so every Gemini run in the launch sweep was quietly at the default. Wiring it up is why Gemini 3 Flash finally has an effort ladder in the matrix above - and why Gemini 3.6 Flash could be measured properly at all.

Try it yourself #

The method is deliberately reproducible. The prompt, ready to paste into any model:

Create an SVG 3D model of a MacBook Pro 16 in a dark color. Respond with ONLY the SVG markup, no explanation.

Or swap in any object your audience knows intimately - a Coke can, a Vespa, a Stratocaster. The only requirements are that the ground truth lives in your reader's head and that style cannot hide broken geometry. We re-run this gallery when new frontier models ship; the pelican had a good run, and the MacBook has years of discrimination left in it.


Method notes: runs on 2026-07-13/14, 2026-07-21/22 and 2026-07-24 via each provider's public API, sampling left at provider defaults, reasoning effort set through each provider's native parameter (Anthropic output_config.effort, OpenAI reasoning.effort, xAI reasoning_effort, Google thinking_level, Moonshot reasoning_effort). Output capped at 16,000 tokens in the first sweep and 64K thereafter; the cap is stated wherever it changed a result. Exact model ids: claude-fable-5, claude-opus-5, claude-opus-4-8, claude-opus-4-6, claude-sonnet-5, claude-sonnet-4-6, gpt-5.6-sol, gpt-5.6-terra, grok-4.5, gemini-3-flash-preview, gemini-3.6-flash and kimi-k3. Costs come from the providers' own usage-reported token counts at list prices; total measured spend across every run shown here is $13.80. Drawings are unedited; failures are reported as failures rather than re-rolled. The measurements are ours; the prose was drafted with AI assistance and edited by a human.

Playcode keeps all of these models one click apart, with the same effort dial we benchmarked here - so you can run your own MacBook test on a real project instead of taking our word for it. Try it at playcode.io.