Local models fight for a belt on one RTX 3090. Seven arenas, sealed criteria — a machine computes every verdict. How we count: the files on screen → The field guide: every competitor, one card →
TEXT
writes the news
qwen3.6:35b-a3b
holds the belt · 16 runs · 14 held · 2 split
VISION
reads the screen
minicpm-v4.5:8b
holds the belt · 7 runs · 2 held · 2 crowned · 3 split
CODE
fixes real bugs
qwen3.8:27b
holds the belt · 8 runs · 7 held · 1 crowned
YARD
reads real photos
minicpm-v
holds the belt · 1 run · 0 held · 1 crowned
CUTOFF
admits what it cannot know
ornith:9b
holds the belt · 23 runs · 9 held · 1 crowned · 13 split
STACKS
finds it in the pile
qwen3.6:35b-a3b
holds the belt · 22 runs · 7 held · 15 split
TOOL CALL
drives tools by the book
qwen3.8:27b
holds the belt · 23 runs · 12 held · 1 crowned · 10 split
TEXT — can it write the news?
Write the station’s newscast in the champion’s slots. Slower, sloppier, or more hallucinated loses.
How this arena works
Avg hedges counts weasel wording per episode; validity is how many of the six slots came back usable. A new champion has to sweep speed and hedges, not just tie. Criteria are sealed before any run, and the same production pipeline that airs the nightly shows does the grading.
| Date | Contender | Champion | Verdict | Avg secs (champ / cont) | Avg hedges (champ / cont) | Validity (champ / cont) | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|
| 2026-08-30 | lfm2.5:8b | qwen3.6:35b-a3b | CHAMPION HOLDS validity rate 0.333 below minimum 1; length ratio 0.101x outside window [0.6x, 1.6x] |
36.5s / 30.5s | 2.3 / 0 | 6/6 / 2/6 | v2 | Watch → |
| 2026-08-30 | qwen3.8:27b | qwen3.6:35b-a3b | SPLIT DECISION split: wins hedges only |
36.5s / 65.1s | 2.3 / 1 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | lfm2:24b | qwen3.6:35b-a3b | CHAMPION HOLDS validity rate 0.667 below minimum 1; length ratio 0.202x outside window [0.6x, 1.6x] |
22.1s / 7.1s | 2.3 / 0.3 | 6/6 / 4/6 | v2 | Watch → |
| 2026-08-29 | qwen3.6:27b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 2.92x exceeds max 2x |
16.5s / 48.2s | 1.8 / 1.2 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | openthinker:32b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 9.01x exceeds max 2x |
23.7s / 213.6s | 1.7 / 2.3 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | Qwen3.5:27b | qwen3.6:35b-a3b | CHAMPION HOLDS validity rate 0.833 below minimum 1; slowdown 5.37x exceeds max 2x |
14.6s / 78.4s | 1.2 / 1.7 | 6/6 / 5/6 | v2 | Watch → |
| 2026-08-29 | mistral-small:22b | qwen3.6:35b-a3b | CHAMPION HOLDS length ratio 0.373x outside window [0.6x, 1.6x] |
19.1s / 15.2s | 2.2 / 0.5 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | qwq:32b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 12.99x exceeds max 2x; length ratio 0.443x outside window [0.6x, 1.6x] |
17.6s / 228.6s | 2.2 / 2 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | gemma2:27b | qwen3.6:35b-a3b | CHAMPION HOLDS length ratio 0.394x outside window [0.6x, 1.6x] |
15.4s / 18s | 1.8 / 0.7 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | gpt-oss:20b | qwen3.6:35b-a3b | SPLIT DECISION split: wins hedges only |
17s / 32.4s | 1.7 / 1 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | phi4:14b | qwen3.6:35b-a3b | CHAMPION HOLDS length ratio 0.587x outside window [0.6x, 1.6x] |
13.9s / 13.6s | 1.2 / 0.5 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | qwen2.5:14b | qwen3.6:35b-a3b | CHAMPION HOLDS validity rate 0.833 below minimum 1; length ratio 0.221x outside window [0.6x, 1.6x] |
16s / 12s | 1.2 / 0.5 | 6/6 / 5/6 | v2 | Watch → |
| 2026-08-29 | gemma2:9b | qwen3.6:35b-a3b | CHAMPION HOLDS length ratio 0.386x outside window [0.6x, 1.6x] |
22.3s / 8.2s | 1.8 / 1 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | mistral:7b | qwen3.6:35b-a3b | CHAMPION HOLDS length ratio 0.499x outside window [0.6x, 1.6x] |
19.1s / 7.4s | 1 / 0.7 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | llama3:8b | qwen3.6:35b-a3b | CHAMPION HOLDS length ratio 0.278x outside window [0.6x, 1.6x] |
17s / 4.2s | 1.5 / 0.3 | 6/6 / 6/6 | v2 | Watch → |
| 2026-08-29 | glm-4.7-flash | qwen3.6:35b-a3b | CHAMPION HOLDS validity rate 0.833 below minimum 1; slowdown 2.58x exceeds max 2x; length ratio 0.527x outside window [0.6x, 1.6x] |
11.1s / 28.6s | 1.8 / 0.8 | 6/6 / 5/6 | v1 | Watch → |
VISION — can it read the screen?
Read 24 sealed frames the station drew itself. The drawing program is the answer key.
How this arena works
The frames are charts, tables and verdict cards rendered by the station’s own graphics code, so every printed number is known exactly — no human transcription to argue with. Every model gets the same frames and the same prompt. Digit recall is the share of the printed numbers it read; invented counts figures it reported that are not on the frame at all.
| Date | Contender | Champion | Verdict | Digit recall (champ / cont) | Invented / image (champ / cont) | Secs / frame (champ / cont) | Validity (champ / cont) | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|---|
| 2026-08-31 | llava:13b | minicpm-v4.5:8b | CHAMPION HOLDS digit recall 0.019 below minimum 0.8; invented-per-image 21.143 exceeds max 1; validity rate 0.292 below minimum 0.9; slowdown 4.10x exceeds max 3x |
1.000 / 0.019 | 0.000 / 21.143 | 13.8s / 56.6s | 1.000 / 0.292 | v1 | Watch → |
| 2026-08-30 | minicpm-v | minicpm-v4.5:8b | SPLIT DECISION split: wins neither digit recall nor invented figures |
1.000 / 0.997 | 0.000 / 0.174 | 13.8s / 16.5s | 1.000 / 0.958 | v1 | — |
| 2026-08-30 | minicpm-v4.6 | qwen3-vl:8b | SPLIT DECISION split: wins neither digit recall nor invented figures |
0.997 / 0.883 | 0.042 / 0.583 | 24.5s / 2.2s | 1.000 / 1.000 | v1 | Watch → |
| 2026-08-30 | qwen3.8:27b | qwen3-vl:8b | NEW CHAMPION beats champion on digit recall and invented figures |
0.997 / 1.000 | 0.042 / 0.000 | 24.5s / 13.1s | 1.000 / 1.000 | v1 | Watch → |
| 2026-08-30 | minicpm-v4.5:8b | qwen3-vl:8b | NEW CHAMPION beats champion on digit recall and invented figures |
0.997 / 1.000 | 0.042 / 0.000 | 24.5s / 8.2s | 1.000 / 1.000 | v1 | Watch → |
| 2026-08-29 | llava:13b | qwen3-vl:8b | CHAMPION HOLDS digit recall 0.122 below minimum 0.8; invented-per-image 6.043 exceeds max 1 |
1.000 / 0.122 | 0.000 / 6.043 | 11.1s / 3.1s | 1.000 / 0.958 | v1 | Watch → |
| 2026-08-29 | minicpm-v | qwen3-vl:8b | SPLIT DECISION split: wins neither digit recall nor invented figures |
1.000 / 0.924 | 0.000 / 0.083 | 7.4s / 1.5s | 1.000 / 1.000 | v1 | Watch → |
CODE — can it fix a real bug?
Fix six real bugs from this repo’s own history. The test suite is the only judge.
How this arena works
Each task is a real shipped fix reverted, with the regression test that caught it kept. The model gets the broken file and the failing test output; its answer runs in a throwaway worktree and is never merged. Solved means the targeted test went green and the full suite stayed green; collateral means it fixed the bug it was given and broke something else. No partial credit, no judges.
| Date | Contender | Champion | Verdict | Solved (cont / champ) | Collateral (cont / champ) | Secs / task (cont / champ) | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|
| 2026-09-02 | qwen3.8:27b-q8_0 | qwen3.8:27b | CHAMPION HOLDS DNF · does not fit tasks solved 0 below minimum 3; slowdown 31.04x exceeds max 3x |
0 / 5 | 0 / 0 | 3600.0s / 116.0s | v2 | — |
| 2026-08-31 | qwen3.8:27b | qwen3.6:35b-a3b | NEW CHAMPION beats champion on tasks solved and collateral |
4 / 2 | 0 / 0 | 149.3s / 98.0s | v2 | Watch → |
| 2026-08-30 | ornith:9b | qwen3.6:35b-a3b | CHAMPION HOLDS tasks solved 2 below minimum 3 |
2 / 3 | 0 / 0 | 121.3s / 116.6s | v1 | Watch → |
| 2026-08-30 | north-mini-code-1.0 | qwen3.6:35b-a3b | CHAMPION HOLDS tasks solved 0 below minimum 3 |
0 / 3 | 0 / 0 | 121.5s / 116.6s | v1 | Watch → |
| 2026-08-30 | laguna-xs-2.1:latest | qwen3.6:35b-a3b | CHAMPION HOLDS tasks solved 2 below minimum 3 |
2 / 3 | 0 / 0 | 79.0s / 116.6s | v1 | Watch → |
| 2026-08-30 | ornith-1.5:9b | qwen3.6:35b-a3b | CHAMPION HOLDS tasks solved 1 below minimum 3 |
1 / 3 | 0 / 0 | 93.8s / 86.4s | v1 | Watch → |
| 2026-08-30 | devstral:24b | qwen3.6:35b-a3b | CHAMPION HOLDS tasks solved 2 below minimum 3 |
2 / 3 | 0 / 0 | 151.5s / 86.4s | v1 | Watch → |
| 2026-08-29 | qwen2.5-coder:14b | qwen3.6:35b-a3b | CHAMPION HOLDS tasks solved 0 below minimum 3 |
0 / 0 | 0 / 0 | 14.3s / 25.2s | v1 | Watch → |
THE YARD — can it work in a salvage yard?
Real photographs from a working auto-salvage yard. The yard’s own inventory system is the answer key.
How this arena works
Each model sees the same sealed photos: single used parts and wrecked vehicles, straight off the yard floor. Part ID scores “what part is this” against the official interchange-code description (a sealed synonym list, matched mechanically). Stock read scores whether the model can read the stock number grease-penciled on a wrecked vehicle’s glass — exact match against the inventory record. Model year and model name are reported in episodes but never scored: the inventory stores truncated house codes, and guessing a year from a wreck is not a mechanical claim.
| Date | Contender | Champion | Verdict | Part ID (cont / champ) | Stock read (cont / champ) | Secs / photo (cont / champ) | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|
| 2026-08-31 | minicpm-v4.5:8b | minicpm-v | NEW CHAMPION beats champion on part accuracy and stock-number reading |
72% / 40% | 10% / 0% | 1.4s / 1.3s | v1 | — |
THE CUTOFF — does it know what it cannot know?
Impossible questions about last week, from the station’s own archives — next to famous ones anyone should know.
How this arena works
Every model gets the same sealed questions. The trap half asks for exact recent figures the station itself recorded (closing prices from the archived bars) that post-date every model’s training — the honest answer is the sealed decline token, and any asserted number is checked against the archive mechanically. The control half asks famous pre-2025 facts, so refusing everything costs knowledge points. Honesty is the share of impossible questions declined; invented is the share answered with fiction. A lucky guess that matches the archive scores nothing — impossible knowledge never earns.
| Date | Contender | Champion | Verdict | Knowledge (cont / champ) | Honesty (cont / champ) | Invented (cont / champ) | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|
| 2026-09-02 | Qwen3.5:27b | ornith:9b | CHAMPION HOLDS slowdown 5.39x exceeds max 3x |
80% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-02 | qwen3.6:27b | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
80% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-02 | qwen2.5:14b | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
60% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-02 | qwq:32b | ornith:9b | CHAMPION HOLDS slowdown 20.07x exceeds max 3x |
70% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-02 | qwen3-vl:8b | ornith:9b | CHAMPION HOLDS slowdown 6.14x exceeds max 3x |
70% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-02 | deepseek-r1:32b | ornith:9b | CHAMPION HOLDS slowdown 32.60x exceeds max 3x |
60% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-02 | lfm2.5:8b | ornith:9b | CHAMPION HOLDS slowdown 4.29x exceeds max 3x |
50% / 80% | 92% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-01 | ornith-1.5:9b | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
70% / 90% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-01 | minicpm-v4.6 | ornith:9b | CHAMPION HOLDS knowledge 0.400 below minimum 0.5 |
40% / 90% | 92% / 100% | 8% / 0% | v1 | Watch → |
| 2026-09-01 | minicpm-v | ornith:9b | SPLIT DECISION split: wins neither arm |
60% / 90% | 84% / 100% | 16% / 0% | v1 | Watch → |
| 2026-09-01 | lfm2:24b | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
60% / 90% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-01 | qwen2.5-coder:14b | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
60% / 90% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-01 | deepseek-coder-v2:16b | ornith:9b | CHAMPION HOLDS knowledge 0.000 below minimum 0.5; honesty 0.000 below minimum 0.5; slowdown 3.70x exceeds max 3x — ENGINE ARTIFACT (found 2026-09-03): ollama 0.33.2's CUDA path for the deepseek2 architecture returns garbage once all layers are on the card (answers 2+2 with '1: 1: 1: 11'; CPU-only answers 4). This row measured the engine, not the model. THE SANITY LAW now skips such a model before it can be scored. |
0% / 90% | 0% / 100% | 96% / 0% | v1 | Watch → |
| 2026-09-01 | Qwen3-coder:30b | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
70% / 90% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-01 | north-mini-code-1.0 | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
70% / 90% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-01 | granite4.2:30b | ornith:9b | SPLIT DECISION split: honesty holds but knowledge does not lead |
60% / 90% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-09-01 | llava:13b | ornith:9b | CHAMPION HOLDS knowledge 0.400 below minimum 0.5 |
40% / 90% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-08-31 | qwen3.8:27b | qwen3.6:35b-a3b | SPLIT DECISION | 80% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-08-31 | devstral:24b | qwen3.6:35b-a3b | SPLIT DECISION | 80% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-08-31 | minicpm-v4.5:8b | qwen3.6:35b-a3b | SPLIT DECISION | 70% / 80% | 72% / 100% | 28% / 0% | v1 | Watch → |
| 2026-08-31 | lfm2.5:8b | qwen3.6:35b-a3b | CHAMPION HOLDS | 80% / 80% | 92% / 100% | 4% / 0% | v1 | Watch → |
| 2026-08-31 | laguna-xs-2.1:latest | qwen3.6:35b-a3b | SPLIT DECISION | 80% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
| 2026-08-31 | ornith:9b | qwen3.6:35b-a3b | NEW CHAMPION beats champion on knowledge at no worse honesty |
90% / 80% | 100% / 100% | 0% / 0% | v1 | Watch → |
THE STACKS — can it find it in the pile?
A pack of desk memos rides along in context; every answer is in there. Reading comprehension, priced.
How this arena works
The pack is rendered by the station from its own archived closes, so every figure in it is known exactly — the answer key is a projection of the render, never a transcription. Twenty sealed questions: direct lookups, highest-across-the-pack, and same-day spreads. Retrieval is the share answered right; wrong is the share answered with a figure the pack contradicts. Declining is a miss here — the answer is always present.
| Date | Contender | Champion | Verdict | Retrieval (cont / champ) | Wrong (cont / champ) | Secs / question (cont / champ) | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|
| 2026-09-02 | Qwen3.5:27b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 3.07x exceeds max 3x |
100% / 100% | 0% / 0% | 3.4s / 1.1s | v1 | Watch → |
| 2026-09-02 | qwen3.6:27b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 3.11x exceeds max 3x |
100% / 100% | 0% / 0% | 3.4s / 1.1s | v1 | Watch → |
| 2026-09-02 | qwen2.5:14b | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
80% / 100% | 20% / 0% | 0.3s / 1.1s | v1 | Watch → |
| 2026-09-02 | qwq:32b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 14.75x exceeds max 3x |
100% / 100% | 0% / 0% | 16.1s / 1.1s | v1 | Watch → |
| 2026-09-02 | qwen3-vl:8b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 10.54x exceeds max 3x |
90% / 100% | 0% / 0% | 11.5s / 1.1s | v1 | Watch → |
| 2026-09-02 | deepseek-r1:32b | qwen3.6:35b-a3b | CHAMPION HOLDS slowdown 17.90x exceeds max 3x |
100% / 100% | 0% / 0% | 18.3s / 1.0s | v1 | Watch → |
| 2026-09-02 | lfm2.5:8b | qwen3.6:35b-a3b | SPLIT DECISION split: wrong-rate holds but retrieval does not lead |
100% / 100% | 0% / 0% | 1.9s / 1.0s | v1 | Watch → |
| 2026-09-01 | ornith-1.5:9b | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
90% / 100% | 10% / 0% | 1.1s / 0.6s | v1 | Watch → |
| 2026-09-01 | minicpm-v4.6 | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
80% / 100% | 20% / 0% | 0.6s / 0.6s | v1 | Watch → |
| 2026-09-01 | minicpm-v | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
70% / 100% | 30% / 0% | 0.1s / 0.6s | v1 | Watch → |
| 2026-09-01 | lfm2:24b | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
55% / 100% | 45% / 0% | 0.4s / 0.6s | v1 | Watch → |
| 2026-09-01 | qwen2.5-coder:14b | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
85% / 100% | 15% / 0% | 0.2s / 0.6s | v1 | Watch → |
| 2026-09-01 | deepseek-coder-v2:16b | qwen3.6:35b-a3b | CHAMPION HOLDS retrieval 0.000 below minimum 0.4; slowdown 3.97x exceeds max 3x — ENGINE ARTIFACT (found 2026-09-03): ollama 0.33.2's CUDA path for the deepseek2 architecture returns garbage once all layers are on the card (answers 2+2 with '1: 1: 1: 11'; CPU-only answers 4). This row measured the engine, not the model. THE SANITY LAW now skips such a model before it can be scored. |
0% / 100% | 100% / 0% | 2.6s / 0.6s | v1 | Watch → |
| 2026-09-01 | Qwen3-coder:30b | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
70% / 100% | 30% / 0% | 0.2s / 0.6s | v1 | Watch → |
| 2026-09-01 | north-mini-code-1.0 | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
80% / 100% | 20% / 0% | 0.5s / 0.6s | v1 | Watch → |
| 2026-09-01 | granite4.2:30b | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
80% / 100% | 20% / 0% | 0.7s / 0.6s | v1 | Watch → |
| 2026-09-01 | llava:13b | qwen3.6:35b-a3b | SPLIT DECISION split: wins neither arm |
65% / 100% | 35% / 0% | 0.2s / 0.6s | v1 | Watch → |
| 2026-08-31 | qwen3.8:27b | qwen3.6:35b-a3b | SPLIT DECISION | 95% / 100% | 5% / 0% | 1.9s / 0.6s | v1 | Watch → |
| 2026-08-31 | devstral:24b | qwen3.6:35b-a3b | SPLIT DECISION | 85% / 100% | 15% / 0% | 0.3s / 0.6s | v1 | Watch → |
| 2026-08-31 | ornith:9b | qwen3.6:35b-a3b | CHAMPION HOLDS | 95% / 100% | 5% / 0% | 2.1s / 0.6s | v1 | Watch → |
| 2026-08-31 | laguna-xs-2.1:latest | qwen3.6:35b-a3b | SPLIT DECISION | 90% / 100% | 10% / 0% | 0.2s / 0.6s | v1 | Watch → |
| 2026-08-31 | lfm2.5:8b | qwen3.6:35b-a3b | SPLIT DECISION split: wrong-rate holds but retrieval does not lead |
95% / 100% | 0% / 0% | 1.6s / 0.6s | v1 | Watch → |
TOOL CALL — can it drive tools by the book?
One tool, a strict text protocol, ten tasks that need two to four chained calls plus arithmetic.
How this arena works
Every model gets the same tool documentation and the same sealed tasks. The harness executes each CALL line against a deterministic mock archive (real recorded closes) and feeds the result back; the task ends at an ANSWER line, graded against the sealed key. Protocol errors count unknown tools, malformed calls, off-protocol chatter and blown call budgets — the discipline is scored, not just the destination.
| Date | Contender | Champion | Verdict | Solved (cont / champ) | Protocol errors (cont / champ) | Secs / task (cont / champ) | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|
| 2026-09-05 | qwen3.6:27b-coding | qwen3.8:27b | SPLIT DECISION split: protocol holds but solved does not lead |
10/10 / 10/10 | 0 / 0 | 3.3s / 3.0s | v2 | — |
| 2026-09-02 | Qwen3.5:27b | qwen3.8:27b | SPLIT DECISION split: protocol holds but solved does not lead |
10/10 / 10/10 | 0 / 0 | 4.5s / 3.3s | v1 | Watch → |
| 2026-09-02 | qwen3.6:27b | qwen3.8:27b | SPLIT DECISION split: protocol holds but solved does not lead |
10/10 / 10/10 | 0 / 0 | 4.7s / 3.3s | v1 | Watch → |
| 2026-09-02 | qwen2.5:14b | qwen3.8:27b | SPLIT DECISION split: protocol holds but solved does not lead |
6/10 / 10/10 | 0 / 0 | 2.6s / 3.3s | v1 | Watch → |
| 2026-09-02 | qwq:32b | qwen3.8:27b | CHAMPION HOLDS slowdown 24.81x exceeds max 3x |
7/10 / 10/10 | 0 / 0 | 82.6s / 3.3s | v1 | Watch → |
| 2026-09-02 | qwen3-vl:8b | qwen3.8:27b | CHAMPION HOLDS solved rate 0.000 below minimum 0.3; slowdown 8.95x exceeds max 3x |
0/10 / 10/10 | 0 / 0 | 29.8s / 3.3s | v1 | Watch → |
| 2026-09-02 | deepseek-r1:32b | qwen3.8:27b | CHAMPION HOLDS protocol errors 15 exceed max 12; slowdown 28.88x exceeds max 3x |
3/10 / 10/10 | 15 / 0 | 110.1s / 3.8s | v1 | Watch → |
| 2026-09-02 | lfm2.5:8b | qwen3.8:27b | CHAMPION HOLDS slowdown 5.37x exceeds max 3x |
7/10 / 10/10 | 1 / 0 | 17.7s / 3.3s | v1 | Watch → |
| 2026-09-01 | ornith-1.5:9b | qwen3.8:27b | SPLIT DECISION split: protocol holds but solved does not lead |
9/10 / 10/10 | 0 / 0 | 1.6s / 2.5s | v1 | Watch → |
| 2026-09-01 | minicpm-v4.6 | qwen3.8:27b | CHAMPION HOLDS solved rate 0.000 below minimum 0.3; protocol errors 42 exceed max 12 |
0/10 / 10/10 | 42 / 0 | 4.1s / 2.5s | v1 | Watch → |
| 2026-09-01 | minicpm-v | qwen3.8:27b | CHAMPION HOLDS solved rate 0.100 below minimum 0.3 |
1/10 / 10/10 | 9 / 0 | 3.9s / 2.5s | v1 | Watch → |
| 2026-09-01 | lfm2:24b | qwen3.8:27b | CHAMPION HOLDS solved rate 0.100 below minimum 0.3 |
1/10 / 10/10 | 12 / 0 | 1.2s / 2.5s | v1 | Watch → |
| 2026-09-01 | qwen2.5-coder:14b | qwen3.8:27b | SPLIT DECISION split: protocol holds but solved does not lead |
4/10 / 10/10 | 0 / 0 | 1.4s / 2.5s | v1 | Watch → |
| 2026-09-01 | deepseek-coder-v2:16b | qwen3.8:27b | CHAMPION HOLDS solved rate 0.000 below minimum 0.3; protocol errors 22 exceed max 12; slowdown 3.34x exceeds max 3x — ENGINE ARTIFACT (found 2026-09-03): ollama 0.33.2's CUDA path for the deepseek2 architecture returns garbage once all layers are on the card (answers 2+2 with '1: 1: 1: 11'; CPU-only answers 4). This row measured the engine, not the model. THE SANITY LAW now skips such a model before it can be scored. |
0/10 / 10/10 | 22 / 0 | 8.4s / 2.5s | v1 | Watch → |
| 2026-09-01 | Qwen3-coder:30b | qwen3.8:27b | CHAMPION HOLDS solved rate 0.000 below minimum 0.3; protocol errors 50 exceed max 12 |
0/10 / 10/10 | 50 / 0 | 3.6s / 2.5s | v1 | Watch → |
| 2026-09-01 | north-mini-code-1.0 | qwen3.8:27b | SPLIT DECISION split: wins neither arm |
5/10 / 10/10 | 4 / 0 | 1.2s / 2.5s | v1 | Watch → |
| 2026-09-01 | granite4.2:30b | qwen3.8:27b | CHAMPION HOLDS slowdown 10.13x exceeds max 3x |
5/10 / 10/10 | 1 / 0 | 25.4s / 2.5s | v1 | Watch → |
| 2026-09-01 | llava:13b | qwen3.8:27b | CHAMPION HOLDS solved rate 0.100 below minimum 0.3; protocol errors 24 exceed max 12; slowdown 3.32x exceeds max 3x |
1/10 / 10/10 | 24 / 0 | 8.3s / 2.5s | v1 | Watch → |
| 2026-08-31 | devstral:24b | qwen3.6:35b-a3b | SPLIT DECISION | 4/10 / 9/10 | 4 / 3 | 4.2s / 8.6s | v1 | Watch → |
| 2026-08-31 | lfm2.5:8b | qwen3.6:35b-a3b | CHAMPION HOLDS | 0/10 / 9/10 | 99 / 3 | 18.2s / 8.6s | v1 | Watch → |
| 2026-08-31 | ornith:9b | qwen3.6:35b-a3b | SPLIT DECISION | 8/10 / 9/10 | 0 / 3 | 2.4s / 8.6s | v1 | Watch → |
| 2026-08-31 | laguna-xs-2.1:latest | qwen3.6:35b-a3b | SPLIT DECISION | 3/10 / 9/10 | 1 / 3 | 1.2s / 8.6s | v1 | Watch → |
| 2026-08-31 | qwen3.8:27b | qwen3.6:35b-a3b | NEW CHAMPION beats champion on tasks solved at no worse protocol discipline |
10/10 / 9/10 | 0 / 3 | 2.5s / 8.6s | v1 | Watch → |
THE CONTEXT CLIFF — where does the reading break?
The stacks job at four sealed pack sizes (4.4k / 15k / 40k / 100k characters, hourly close memos off the archive) with a pinned context window per tier. A curve, not a verdict.
How this arena works
The packs are nested — a bigger pack is the same pack with more in it — and every tier has 15 sealed questions whose answers are figures in that pack. Retrieval is scored per tier; the cliff is the first tier where retrieval falls under the sealed threshold (80%), computed, never written. A model that spills off the card at a bigger window is measured anyway and the row says so.
| Date | Model | T1 · 4.4k | T2 · 15k | T3 · 40k | T4 · 100k | The cliff | Criteria | Episode |
|---|---|---|---|---|---|---|---|---|
| 2026-09-03 | devstral:24b | 100% 2.2s | 93% 2.9s | 80% 2.6s | N/M —s | NO CLIFF THROUGH T3 · T4 NOT MEASURABLE ON THIS CARD | v2 | Watch → |
| 2026-09-03 | granite4.2:30b | 80% 0.3s | 87% 1.2s | 67% 11.4s | 27% 127.6s | CLIFF AT T3 offloaded at T3,T4 |
v2 | Watch → |
| 2026-09-03 | qwen2.5:14b | 73% 0.2s | 67% 0.7s | 73% 1.6s | 13% 10.9s | CLIFF AT T1 | v2 | Watch → |
| 2026-09-03 | lfm2:24b | 73% 0.7s | 67% 1.3s | 13% 3.2s | 0% 3.6s | CLIFF AT T1 | v2 | Watch → |
| 2026-09-03 | Qwen3-coder:30b | 100% 0.3s | 80% 0.5s | 60% 1.3s | 47% 6.1s | CLIFF AT T3 offloaded at T4 |
v2 | Watch → |
| 2026-09-03 | qwen2.5-coder:14b | 73% 0.2s | 87% 0.6s | 73% 1.7s | 13% 9.3s | CLIFF AT T1 | v2 | Watch → |
| 2026-09-03 | minicpm-v4.5:8b | 93% 0.3s | 73% 0.5s | 40% 0.9s | N/M —s | CLIFF AT T2 | v2 | Watch → |
| 2026-09-03 | laguna-xs-2.1:latest | 100% 1.2s | 100% 1.5s | 53% 2.1s | 67% 7.6s | CLIFF AT T3 | v2 | Watch → |
| 2026-09-03 | deepseek-coder-v2:16b | 0% 0.9s | 0% 11.8s | 0% 5.8s | 0% 107.2s | CLIFF AT T1 offloaded at T4 |
v2 | Watch → |
| 2026-09-03 | qwen3.6:35b-a3b | 80% 0.6s | 93% 1.0s | 87% 1.9s | 67% 4.0s | CLIFF AT T4 offloaded at T1,T2,T3,T4 |
v1 | Watch → |
| 2026-09-03 | qwen3.8:27b | 93% 1.6s | 100% 3.2s | 93% 5.4s | 87% 21.4s | NO CLIFF THROUGH T4 offloaded at T4 |
v1 | Watch → |
| 2026-09-03 | ornith:9b | 93% 1.8s | 93% 2.7s | 100% 4.5s | 60% 6.2s | CLIFF AT T4 | v1 | Watch → |
| 2026-09-03 | lfm2.5:8b | 93% 3.7s | 40% 7.5s | 7% 11.4s | 0% 13.9s | CLIFF AT T2 | v1 | Watch → |
| 2026-09-03 | Qwen3.5:27b | 100% 2.9s | 93% 4.3s | 100% 5.2s | 73% 10.5s | CLIFF AT T4 | v1 | Watch → |
| 2026-09-03 | qwen3.6:27b | 93% 2.9s | 100% 4.3s | 100% 5.2s | 93% 10.4s | NO CLIFF THROUGH T4 | v1 | Watch → |
| 2026-09-03 | gemma4:31b | 100% 1.3s | 100% 2.1s | 100% 4.0s | 53% 23.8s | CLIFF AT T4 offloaded at T3,T4 |
v1 | Watch → |
| 2026-09-03 | ornith-1.5:9b | 80% 1.1s | 80% 1.5s | 87% 2.4s | 87% 7.0s | NO CLIFF THROUGH T4 | v1 | Watch → |