THE GAUNTLET — local LLM ledger, one RTX 3090

20 min read Original article ↗
← INFORMANT

Local models fight for a belt on one RTX 3090. Seven arenas, sealed criteria — a machine computes every verdict. How we count: the files on screen → The field guide: every competitor, one card →

TEXT

writes the news

qwen3.6:35b-a3b

holds the belt · 16 runs · 14 held · 2 split

VISION

reads the screen

minicpm-v4.5:8b

holds the belt · 7 runs · 2 held · 2 crowned · 3 split

CODE

fixes real bugs

qwen3.8:27b

holds the belt · 8 runs · 7 held · 1 crowned

YARD

reads real photos

minicpm-v

holds the belt · 1 run · 0 held · 1 crowned

CUTOFF

admits what it cannot know

ornith:9b

holds the belt · 23 runs · 9 held · 1 crowned · 13 split

STACKS

finds it in the pile

qwen3.6:35b-a3b

holds the belt · 22 runs · 7 held · 15 split

TOOL CALL

drives tools by the book

qwen3.8:27b

holds the belt · 23 runs · 12 held · 1 crowned · 10 split

TEXT — can it write the news?

Write the station’s newscast in the champion’s slots. Slower, sloppier, or more hallucinated loses.

How this arena works

Avg hedges counts weasel wording per episode; validity is how many of the six slots came back usable. A new champion has to sweep speed and hedges, not just tie. Criteria are sealed before any run, and the same production pipeline that airs the nightly shows does the grading.

Ledger of every bake-off run, newest first.
Date Contender Champion Verdict Avg secs (champ / cont) Avg hedges (champ / cont) Validity (champ / cont) Criteria Episode
2026-08-30 lfm2.5:8b qwen3.6:35b-a3b CHAMPION HOLDS

validity rate 0.333 below minimum 1; length ratio 0.101x outside window [0.6x, 1.6x]

36.5s / 30.5s 2.3 / 0 6/6 / 2/6 v2 Watch →
2026-08-30 qwen3.8:27b qwen3.6:35b-a3b SPLIT DECISION

split: wins hedges only

36.5s / 65.1s 2.3 / 1 6/6 / 6/6 v2 Watch →
2026-08-29 lfm2:24b qwen3.6:35b-a3b CHAMPION HOLDS

validity rate 0.667 below minimum 1; length ratio 0.202x outside window [0.6x, 1.6x]

22.1s / 7.1s 2.3 / 0.3 6/6 / 4/6 v2 Watch →
2026-08-29 qwen3.6:27b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 2.92x exceeds max 2x

16.5s / 48.2s 1.8 / 1.2 6/6 / 6/6 v2 Watch →
2026-08-29 openthinker:32b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 9.01x exceeds max 2x

23.7s / 213.6s 1.7 / 2.3 6/6 / 6/6 v2 Watch →
2026-08-29 Qwen3.5:27b qwen3.6:35b-a3b CHAMPION HOLDS

validity rate 0.833 below minimum 1; slowdown 5.37x exceeds max 2x

14.6s / 78.4s 1.2 / 1.7 6/6 / 5/6 v2 Watch →
2026-08-29 mistral-small:22b qwen3.6:35b-a3b CHAMPION HOLDS

length ratio 0.373x outside window [0.6x, 1.6x]

19.1s / 15.2s 2.2 / 0.5 6/6 / 6/6 v2 Watch →
2026-08-29 qwq:32b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 12.99x exceeds max 2x; length ratio 0.443x outside window [0.6x, 1.6x]

17.6s / 228.6s 2.2 / 2 6/6 / 6/6 v2 Watch →
2026-08-29 gemma2:27b qwen3.6:35b-a3b CHAMPION HOLDS

length ratio 0.394x outside window [0.6x, 1.6x]

15.4s / 18s 1.8 / 0.7 6/6 / 6/6 v2 Watch →
2026-08-29 gpt-oss:20b qwen3.6:35b-a3b SPLIT DECISION

split: wins hedges only

17s / 32.4s 1.7 / 1 6/6 / 6/6 v2 Watch →
2026-08-29 phi4:14b qwen3.6:35b-a3b CHAMPION HOLDS

length ratio 0.587x outside window [0.6x, 1.6x]

13.9s / 13.6s 1.2 / 0.5 6/6 / 6/6 v2 Watch →
2026-08-29 qwen2.5:14b qwen3.6:35b-a3b CHAMPION HOLDS

validity rate 0.833 below minimum 1; length ratio 0.221x outside window [0.6x, 1.6x]

16s / 12s 1.2 / 0.5 6/6 / 5/6 v2 Watch →
2026-08-29 gemma2:9b qwen3.6:35b-a3b CHAMPION HOLDS

length ratio 0.386x outside window [0.6x, 1.6x]

22.3s / 8.2s 1.8 / 1 6/6 / 6/6 v2 Watch →
2026-08-29 mistral:7b qwen3.6:35b-a3b CHAMPION HOLDS

length ratio 0.499x outside window [0.6x, 1.6x]

19.1s / 7.4s 1 / 0.7 6/6 / 6/6 v2 Watch →
2026-08-29 llama3:8b qwen3.6:35b-a3b CHAMPION HOLDS

length ratio 0.278x outside window [0.6x, 1.6x]

17s / 4.2s 1.5 / 0.3 6/6 / 6/6 v2 Watch →
2026-08-29 glm-4.7-flash qwen3.6:35b-a3b CHAMPION HOLDS

validity rate 0.833 below minimum 1; slowdown 2.58x exceeds max 2x; length ratio 0.527x outside window [0.6x, 1.6x]

11.1s / 28.6s 1.8 / 0.8 6/6 / 5/6 v1 Watch →

VISION — can it read the screen?

Read 24 sealed frames the station drew itself. The drawing program is the answer key.

How this arena works

The frames are charts, tables and verdict cards rendered by the station’s own graphics code, so every printed number is known exactly — no human transcription to argue with. Every model gets the same frames and the same prompt. Digit recall is the share of the printed numbers it read; invented counts figures it reported that are not on the frame at all.

Vision ledger: every model that has read the sealed frames, newest first.
Date Contender Champion Verdict Digit recall (champ / cont) Invented / image (champ / cont) Secs / frame (champ / cont) Validity (champ / cont) Criteria Episode
2026-08-31 llava:13b minicpm-v4.5:8b CHAMPION HOLDS

digit recall 0.019 below minimum 0.8; invented-per-image 21.143 exceeds max 1; validity rate 0.292 below minimum 0.9; slowdown 4.10x exceeds max 3x

1.000 / 0.019 0.000 / 21.143 13.8s / 56.6s 1.000 / 0.292 v1 Watch →
2026-08-30 minicpm-v minicpm-v4.5:8b SPLIT DECISION

split: wins neither digit recall nor invented figures

1.000 / 0.997 0.000 / 0.174 13.8s / 16.5s 1.000 / 0.958 v1
2026-08-30 minicpm-v4.6 qwen3-vl:8b SPLIT DECISION

split: wins neither digit recall nor invented figures

0.997 / 0.883 0.042 / 0.583 24.5s / 2.2s 1.000 / 1.000 v1 Watch →
2026-08-30 qwen3.8:27b qwen3-vl:8b NEW CHAMPION

beats champion on digit recall and invented figures

0.997 / 1.000 0.042 / 0.000 24.5s / 13.1s 1.000 / 1.000 v1 Watch →
2026-08-30 minicpm-v4.5:8b qwen3-vl:8b NEW CHAMPION

beats champion on digit recall and invented figures

0.997 / 1.000 0.042 / 0.000 24.5s / 8.2s 1.000 / 1.000 v1 Watch →
2026-08-29 llava:13b qwen3-vl:8b CHAMPION HOLDS

digit recall 0.122 below minimum 0.8; invented-per-image 6.043 exceeds max 1

1.000 / 0.122 0.000 / 6.043 11.1s / 3.1s 1.000 / 0.958 v1 Watch →
2026-08-29 minicpm-v qwen3-vl:8b SPLIT DECISION

split: wins neither digit recall nor invented figures

1.000 / 0.924 0.000 / 0.083 7.4s / 1.5s 1.000 / 1.000 v1 Watch →

CODE — can it fix a real bug?

Fix six real bugs from this repo’s own history. The test suite is the only judge.

How this arena works

Each task is a real shipped fix reverted, with the regression test that caught it kept. The model gets the broken file and the failing test output; its answer runs in a throwaway worktree and is never merged. Solved means the targeted test went green and the full suite stayed green; collateral means it fixed the bug it was given and broke something else. No partial credit, no judges.

Code ledger: every model that has been handed the sealed bugs, newest first.
Date Contender Champion Verdict Solved (cont / champ) Collateral (cont / champ) Secs / task (cont / champ) Criteria Episode
2026-09-02 qwen3.8:27b-q8_0 qwen3.8:27b CHAMPION HOLDS DNF · does not fit

tasks solved 0 below minimum 3; slowdown 31.04x exceeds max 3x

0 / 5 0 / 0 3600.0s / 116.0s v2
2026-08-31 qwen3.8:27b qwen3.6:35b-a3b NEW CHAMPION

beats champion on tasks solved and collateral

4 / 2 0 / 0 149.3s / 98.0s v2 Watch →
2026-08-30 ornith:9b qwen3.6:35b-a3b CHAMPION HOLDS

tasks solved 2 below minimum 3

2 / 3 0 / 0 121.3s / 116.6s v1 Watch →
2026-08-30 north-mini-code-1.0 qwen3.6:35b-a3b CHAMPION HOLDS

tasks solved 0 below minimum 3

0 / 3 0 / 0 121.5s / 116.6s v1 Watch →
2026-08-30 laguna-xs-2.1:latest qwen3.6:35b-a3b CHAMPION HOLDS

tasks solved 2 below minimum 3

2 / 3 0 / 0 79.0s / 116.6s v1 Watch →
2026-08-30 ornith-1.5:9b qwen3.6:35b-a3b CHAMPION HOLDS

tasks solved 1 below minimum 3

1 / 3 0 / 0 93.8s / 86.4s v1 Watch →
2026-08-30 devstral:24b qwen3.6:35b-a3b CHAMPION HOLDS

tasks solved 2 below minimum 3

2 / 3 0 / 0 151.5s / 86.4s v1 Watch →
2026-08-29 qwen2.5-coder:14b qwen3.6:35b-a3b CHAMPION HOLDS

tasks solved 0 below minimum 3

0 / 0 0 / 0 14.3s / 25.2s v1 Watch →

THE YARD — can it work in a salvage yard?

Real photographs from a working auto-salvage yard. The yard’s own inventory system is the answer key.

How this arena works

Each model sees the same sealed photos: single used parts and wrecked vehicles, straight off the yard floor. Part ID scores “what part is this” against the official interchange-code description (a sealed synonym list, matched mechanically). Stock read scores whether the model can read the stock number grease-penciled on a wrecked vehicle’s glass — exact match against the inventory record. Model year and model name are reported in episodes but never scored: the inventory stores truncated house codes, and guessing a year from a wreck is not a mechanical claim.

Yard ledger: every model that has worked the sealed photos, newest first.
Date Contender Champion Verdict Part ID (cont / champ) Stock read (cont / champ) Secs / photo (cont / champ) Criteria Episode
2026-08-31 minicpm-v4.5:8b minicpm-v NEW CHAMPION

beats champion on part accuracy and stock-number reading

72% / 40% 10% / 0% 1.4s / 1.3s v1

THE CUTOFF — does it know what it cannot know?

Impossible questions about last week, from the station’s own archives — next to famous ones anyone should know.

How this arena works

Every model gets the same sealed questions. The trap half asks for exact recent figures the station itself recorded (closing prices from the archived bars) that post-date every model’s training — the honest answer is the sealed decline token, and any asserted number is checked against the archive mechanically. The control half asks famous pre-2025 facts, so refusing everything costs knowledge points. Honesty is the share of impossible questions declined; invented is the share answered with fiction. A lucky guess that matches the archive scores nothing — impossible knowledge never earns.

Cutoff ledger: every model that has faced the impossible questions, newest first.
Date Contender Champion Verdict Knowledge (cont / champ) Honesty (cont / champ) Invented (cont / champ) Criteria Episode
2026-09-02 Qwen3.5:27b ornith:9b CHAMPION HOLDS

slowdown 5.39x exceeds max 3x

80% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-09-02 qwen3.6:27b ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

80% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-09-02 qwen2.5:14b ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

60% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-09-02 qwq:32b ornith:9b CHAMPION HOLDS

slowdown 20.07x exceeds max 3x

70% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-09-02 qwen3-vl:8b ornith:9b CHAMPION HOLDS

slowdown 6.14x exceeds max 3x

70% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-09-02 deepseek-r1:32b ornith:9b CHAMPION HOLDS

slowdown 32.60x exceeds max 3x

60% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-09-02 lfm2.5:8b ornith:9b CHAMPION HOLDS

slowdown 4.29x exceeds max 3x

50% / 80% 92% / 100% 0% / 0% v1 Watch →
2026-09-01 ornith-1.5:9b ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

70% / 90% 100% / 100% 0% / 0% v1 Watch →
2026-09-01 minicpm-v4.6 ornith:9b CHAMPION HOLDS

knowledge 0.400 below minimum 0.5

40% / 90% 92% / 100% 8% / 0% v1 Watch →
2026-09-01 minicpm-v ornith:9b SPLIT DECISION

split: wins neither arm

60% / 90% 84% / 100% 16% / 0% v1 Watch →
2026-09-01 lfm2:24b ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

60% / 90% 100% / 100% 0% / 0% v1 Watch →
2026-09-01 qwen2.5-coder:14b ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

60% / 90% 100% / 100% 0% / 0% v1 Watch →
2026-09-01 deepseek-coder-v2:16b ornith:9b CHAMPION HOLDS

knowledge 0.000 below minimum 0.5; honesty 0.000 below minimum 0.5; slowdown 3.70x exceeds max 3x — ENGINE ARTIFACT (found 2026-09-03): ollama 0.33.2's CUDA path for the deepseek2 architecture returns garbage once all layers are on the card (answers 2+2 with '1: 1: 1: 11'; CPU-only answers 4). This row measured the engine, not the model. THE SANITY LAW now skips such a model before it can be scored.

0% / 90% 0% / 100% 96% / 0% v1 Watch →
2026-09-01 Qwen3-coder:30b ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

70% / 90% 100% / 100% 0% / 0% v1 Watch →
2026-09-01 north-mini-code-1.0 ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

70% / 90% 100% / 100% 0% / 0% v1 Watch →
2026-09-01 granite4.2:30b ornith:9b SPLIT DECISION

split: honesty holds but knowledge does not lead

60% / 90% 100% / 100% 0% / 0% v1 Watch →
2026-09-01 llava:13b ornith:9b CHAMPION HOLDS

knowledge 0.400 below minimum 0.5

40% / 90% 100% / 100% 0% / 0% v1 Watch →
2026-08-31 qwen3.8:27b qwen3.6:35b-a3b SPLIT DECISION 80% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-08-31 devstral:24b qwen3.6:35b-a3b SPLIT DECISION 80% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-08-31 minicpm-v4.5:8b qwen3.6:35b-a3b SPLIT DECISION 70% / 80% 72% / 100% 28% / 0% v1 Watch →
2026-08-31 lfm2.5:8b qwen3.6:35b-a3b CHAMPION HOLDS 80% / 80% 92% / 100% 4% / 0% v1 Watch →
2026-08-31 laguna-xs-2.1:latest qwen3.6:35b-a3b SPLIT DECISION 80% / 80% 100% / 100% 0% / 0% v1 Watch →
2026-08-31 ornith:9b qwen3.6:35b-a3b NEW CHAMPION

beats champion on knowledge at no worse honesty

90% / 80% 100% / 100% 0% / 0% v1 Watch →

THE STACKS — can it find it in the pile?

A pack of desk memos rides along in context; every answer is in there. Reading comprehension, priced.

How this arena works

The pack is rendered by the station from its own archived closes, so every figure in it is known exactly — the answer key is a projection of the render, never a transcription. Twenty sealed questions: direct lookups, highest-across-the-pack, and same-day spreads. Retrieval is the share answered right; wrong is the share answered with a figure the pack contradicts. Declining is a miss here — the answer is always present.

Stacks ledger: every model that has read the pack, newest first.
Date Contender Champion Verdict Retrieval (cont / champ) Wrong (cont / champ) Secs / question (cont / champ) Criteria Episode
2026-09-02 Qwen3.5:27b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 3.07x exceeds max 3x

100% / 100% 0% / 0% 3.4s / 1.1s v1 Watch →
2026-09-02 qwen3.6:27b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 3.11x exceeds max 3x

100% / 100% 0% / 0% 3.4s / 1.1s v1 Watch →
2026-09-02 qwen2.5:14b qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

80% / 100% 20% / 0% 0.3s / 1.1s v1 Watch →
2026-09-02 qwq:32b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 14.75x exceeds max 3x

100% / 100% 0% / 0% 16.1s / 1.1s v1 Watch →
2026-09-02 qwen3-vl:8b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 10.54x exceeds max 3x

90% / 100% 0% / 0% 11.5s / 1.1s v1 Watch →
2026-09-02 deepseek-r1:32b qwen3.6:35b-a3b CHAMPION HOLDS

slowdown 17.90x exceeds max 3x

100% / 100% 0% / 0% 18.3s / 1.0s v1 Watch →
2026-09-02 lfm2.5:8b qwen3.6:35b-a3b SPLIT DECISION

split: wrong-rate holds but retrieval does not lead

100% / 100% 0% / 0% 1.9s / 1.0s v1 Watch →
2026-09-01 ornith-1.5:9b qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

90% / 100% 10% / 0% 1.1s / 0.6s v1 Watch →
2026-09-01 minicpm-v4.6 qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

80% / 100% 20% / 0% 0.6s / 0.6s v1 Watch →
2026-09-01 minicpm-v qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

70% / 100% 30% / 0% 0.1s / 0.6s v1 Watch →
2026-09-01 lfm2:24b qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

55% / 100% 45% / 0% 0.4s / 0.6s v1 Watch →
2026-09-01 qwen2.5-coder:14b qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

85% / 100% 15% / 0% 0.2s / 0.6s v1 Watch →
2026-09-01 deepseek-coder-v2:16b qwen3.6:35b-a3b CHAMPION HOLDS

retrieval 0.000 below minimum 0.4; slowdown 3.97x exceeds max 3x — ENGINE ARTIFACT (found 2026-09-03): ollama 0.33.2's CUDA path for the deepseek2 architecture returns garbage once all layers are on the card (answers 2+2 with '1: 1: 1: 11'; CPU-only answers 4). This row measured the engine, not the model. THE SANITY LAW now skips such a model before it can be scored.

0% / 100% 100% / 0% 2.6s / 0.6s v1 Watch →
2026-09-01 Qwen3-coder:30b qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

70% / 100% 30% / 0% 0.2s / 0.6s v1 Watch →
2026-09-01 north-mini-code-1.0 qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

80% / 100% 20% / 0% 0.5s / 0.6s v1 Watch →
2026-09-01 granite4.2:30b qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

80% / 100% 20% / 0% 0.7s / 0.6s v1 Watch →
2026-09-01 llava:13b qwen3.6:35b-a3b SPLIT DECISION

split: wins neither arm

65% / 100% 35% / 0% 0.2s / 0.6s v1 Watch →
2026-08-31 qwen3.8:27b qwen3.6:35b-a3b SPLIT DECISION 95% / 100% 5% / 0% 1.9s / 0.6s v1 Watch →
2026-08-31 devstral:24b qwen3.6:35b-a3b SPLIT DECISION 85% / 100% 15% / 0% 0.3s / 0.6s v1 Watch →
2026-08-31 ornith:9b qwen3.6:35b-a3b CHAMPION HOLDS 95% / 100% 5% / 0% 2.1s / 0.6s v1 Watch →
2026-08-31 laguna-xs-2.1:latest qwen3.6:35b-a3b SPLIT DECISION 90% / 100% 10% / 0% 0.2s / 0.6s v1 Watch →
2026-08-31 lfm2.5:8b qwen3.6:35b-a3b SPLIT DECISION

split: wrong-rate holds but retrieval does not lead

95% / 100% 0% / 0% 1.6s / 0.6s v1 Watch →

TOOL CALL — can it drive tools by the book?

One tool, a strict text protocol, ten tasks that need two to four chained calls plus arithmetic.

How this arena works

Every model gets the same tool documentation and the same sealed tasks. The harness executes each CALL line against a deterministic mock archive (real recorded closes) and feeds the result back; the task ends at an ANSWER line, graded against the sealed key. Protocol errors count unknown tools, malformed calls, off-protocol chatter and blown call budgets — the discipline is scored, not just the destination.

Tool-call ledger: every model that has driven the archive tool, newest first.
Date Contender Champion Verdict Solved (cont / champ) Protocol errors (cont / champ) Secs / task (cont / champ) Criteria Episode
2026-09-05 qwen3.6:27b-coding qwen3.8:27b SPLIT DECISION

split: protocol holds but solved does not lead

10/10 / 10/10 0 / 0 3.3s / 3.0s v2
2026-09-02 Qwen3.5:27b qwen3.8:27b SPLIT DECISION

split: protocol holds but solved does not lead

10/10 / 10/10 0 / 0 4.5s / 3.3s v1 Watch →
2026-09-02 qwen3.6:27b qwen3.8:27b SPLIT DECISION

split: protocol holds but solved does not lead

10/10 / 10/10 0 / 0 4.7s / 3.3s v1 Watch →
2026-09-02 qwen2.5:14b qwen3.8:27b SPLIT DECISION

split: protocol holds but solved does not lead

6/10 / 10/10 0 / 0 2.6s / 3.3s v1 Watch →
2026-09-02 qwq:32b qwen3.8:27b CHAMPION HOLDS

slowdown 24.81x exceeds max 3x

7/10 / 10/10 0 / 0 82.6s / 3.3s v1 Watch →
2026-09-02 qwen3-vl:8b qwen3.8:27b CHAMPION HOLDS

solved rate 0.000 below minimum 0.3; slowdown 8.95x exceeds max 3x

0/10 / 10/10 0 / 0 29.8s / 3.3s v1 Watch →
2026-09-02 deepseek-r1:32b qwen3.8:27b CHAMPION HOLDS

protocol errors 15 exceed max 12; slowdown 28.88x exceeds max 3x

3/10 / 10/10 15 / 0 110.1s / 3.8s v1 Watch →
2026-09-02 lfm2.5:8b qwen3.8:27b CHAMPION HOLDS

slowdown 5.37x exceeds max 3x

7/10 / 10/10 1 / 0 17.7s / 3.3s v1 Watch →
2026-09-01 ornith-1.5:9b qwen3.8:27b SPLIT DECISION

split: protocol holds but solved does not lead

9/10 / 10/10 0 / 0 1.6s / 2.5s v1 Watch →
2026-09-01 minicpm-v4.6 qwen3.8:27b CHAMPION HOLDS

solved rate 0.000 below minimum 0.3; protocol errors 42 exceed max 12

0/10 / 10/10 42 / 0 4.1s / 2.5s v1 Watch →
2026-09-01 minicpm-v qwen3.8:27b CHAMPION HOLDS

solved rate 0.100 below minimum 0.3

1/10 / 10/10 9 / 0 3.9s / 2.5s v1 Watch →
2026-09-01 lfm2:24b qwen3.8:27b CHAMPION HOLDS

solved rate 0.100 below minimum 0.3

1/10 / 10/10 12 / 0 1.2s / 2.5s v1 Watch →
2026-09-01 qwen2.5-coder:14b qwen3.8:27b SPLIT DECISION

split: protocol holds but solved does not lead

4/10 / 10/10 0 / 0 1.4s / 2.5s v1 Watch →
2026-09-01 deepseek-coder-v2:16b qwen3.8:27b CHAMPION HOLDS

solved rate 0.000 below minimum 0.3; protocol errors 22 exceed max 12; slowdown 3.34x exceeds max 3x — ENGINE ARTIFACT (found 2026-09-03): ollama 0.33.2's CUDA path for the deepseek2 architecture returns garbage once all layers are on the card (answers 2+2 with '1: 1: 1: 11'; CPU-only answers 4). This row measured the engine, not the model. THE SANITY LAW now skips such a model before it can be scored.

0/10 / 10/10 22 / 0 8.4s / 2.5s v1 Watch →
2026-09-01 Qwen3-coder:30b qwen3.8:27b CHAMPION HOLDS

solved rate 0.000 below minimum 0.3; protocol errors 50 exceed max 12

0/10 / 10/10 50 / 0 3.6s / 2.5s v1 Watch →
2026-09-01 north-mini-code-1.0 qwen3.8:27b SPLIT DECISION

split: wins neither arm

5/10 / 10/10 4 / 0 1.2s / 2.5s v1 Watch →
2026-09-01 granite4.2:30b qwen3.8:27b CHAMPION HOLDS

slowdown 10.13x exceeds max 3x

5/10 / 10/10 1 / 0 25.4s / 2.5s v1 Watch →
2026-09-01 llava:13b qwen3.8:27b CHAMPION HOLDS

solved rate 0.100 below minimum 0.3; protocol errors 24 exceed max 12; slowdown 3.32x exceeds max 3x

1/10 / 10/10 24 / 0 8.3s / 2.5s v1 Watch →
2026-08-31 devstral:24b qwen3.6:35b-a3b SPLIT DECISION 4/10 / 9/10 4 / 3 4.2s / 8.6s v1 Watch →
2026-08-31 lfm2.5:8b qwen3.6:35b-a3b CHAMPION HOLDS 0/10 / 9/10 99 / 3 18.2s / 8.6s v1 Watch →
2026-08-31 ornith:9b qwen3.6:35b-a3b SPLIT DECISION 8/10 / 9/10 0 / 3 2.4s / 8.6s v1 Watch →
2026-08-31 laguna-xs-2.1:latest qwen3.6:35b-a3b SPLIT DECISION 3/10 / 9/10 1 / 3 1.2s / 8.6s v1 Watch →
2026-08-31 qwen3.8:27b qwen3.6:35b-a3b NEW CHAMPION

beats champion on tasks solved at no worse protocol discipline

10/10 / 9/10 0 / 3 2.5s / 8.6s v1 Watch →

THE CONTEXT CLIFF — where does the reading break?

The stacks job at four sealed pack sizes (4.4k / 15k / 40k / 100k characters, hourly close memos off the archive) with a pinned context window per tier. A curve, not a verdict.

How this arena works

The packs are nested — a bigger pack is the same pack with more in it — and every tier has 15 sealed questions whose answers are figures in that pack. Retrieval is scored per tier; the cliff is the first tier where retrieval falls under the sealed threshold (80%), computed, never written. A model that spills off the card at a bigger window is measured anyway and the row says so.

Context-cliff ledger: one row per model per run, newest first. Each tier cell is retrieval, then seconds per question.
Date Model T1 · 4.4k T2 · 15k T3 · 40k T4 · 100k The cliff Criteria Episode
2026-09-03 devstral:24b 100% 2.2s 93% 2.9s 80% 2.6s N/M —s NO CLIFF THROUGH T3 · T4 NOT MEASURABLE ON THIS CARD v2 Watch →
2026-09-03 granite4.2:30b 80% 0.3s 87% 1.2s 67% 11.4s 27% 127.6s CLIFF AT T3

offloaded at T3,T4

v2 Watch →
2026-09-03 qwen2.5:14b 73% 0.2s 67% 0.7s 73% 1.6s 13% 10.9s CLIFF AT T1 v2 Watch →
2026-09-03 lfm2:24b 73% 0.7s 67% 1.3s 13% 3.2s 0% 3.6s CLIFF AT T1 v2 Watch →
2026-09-03 Qwen3-coder:30b 100% 0.3s 80% 0.5s 60% 1.3s 47% 6.1s CLIFF AT T3

offloaded at T4

v2 Watch →
2026-09-03 qwen2.5-coder:14b 73% 0.2s 87% 0.6s 73% 1.7s 13% 9.3s CLIFF AT T1 v2 Watch →
2026-09-03 minicpm-v4.5:8b 93% 0.3s 73% 0.5s 40% 0.9s N/M —s CLIFF AT T2 v2 Watch →
2026-09-03 laguna-xs-2.1:latest 100% 1.2s 100% 1.5s 53% 2.1s 67% 7.6s CLIFF AT T3 v2 Watch →
2026-09-03 deepseek-coder-v2:16b 0% 0.9s 0% 11.8s 0% 5.8s 0% 107.2s CLIFF AT T1

offloaded at T4

v2 Watch →
2026-09-03 qwen3.6:35b-a3b 80% 0.6s 93% 1.0s 87% 1.9s 67% 4.0s CLIFF AT T4

offloaded at T1,T2,T3,T4

v1 Watch →
2026-09-03 qwen3.8:27b 93% 1.6s 100% 3.2s 93% 5.4s 87% 21.4s NO CLIFF THROUGH T4

offloaded at T4

v1 Watch →
2026-09-03 ornith:9b 93% 1.8s 93% 2.7s 100% 4.5s 60% 6.2s CLIFF AT T4 v1 Watch →
2026-09-03 lfm2.5:8b 93% 3.7s 40% 7.5s 7% 11.4s 0% 13.9s CLIFF AT T2 v1 Watch →
2026-09-03 Qwen3.5:27b 100% 2.9s 93% 4.3s 100% 5.2s 73% 10.5s CLIFF AT T4 v1 Watch →
2026-09-03 qwen3.6:27b 93% 2.9s 100% 4.3s 100% 5.2s 93% 10.4s NO CLIFF THROUGH T4 v1 Watch →
2026-09-03 gemma4:31b 100% 1.3s 100% 2.1s 100% 4.0s 53% 23.8s CLIFF AT T4

offloaded at T3,T4

v1 Watch →
2026-09-03 ornith-1.5:9b 80% 1.1s 80% 1.5s 87% 2.4s 87% 7.0s NO CLIFF THROUGH T4 v1 Watch →