JevBench v1.3.0: Jev alternatives ranked

Benchmark Heaven

75 min read Original article ↗

JevBench v1.3.0JevBench v1.3.0 · our own benchmark

Jev-class models

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Version 1.3.0 measures 52 systems on the unchanged 534 decisions, including 220 hard ones, and ranks them by the JevBench Score. Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.

Scored 21 Sept 2026 · protocol jevbench::v1.2 · 72 easy + 96 standard + 146 judge + 220 hard decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 20fce8e6e4f0 · v1.0 results

JevBench v1.3.0 · 534 decisions per system

JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)

OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty. Change the weighting ↓

  1. 1Jev 1.13.074.4I 86 · C 83 · S 83 · K 52 · $0.040
  2. 2SemIf (Qwen3.5-4B)73.1I 79 · C 73 · S 84 · K 59 · ~$0.022 est.
  3. 3djev (Maisa, diffusion-gemma)73.0I 83 · C 65 · S 91 · K 58 · $0.026 ann.
  4. 4Winnow-12B Q871.2I 82 · C 72 · S 82 · K 53 · ~$0.037 est.
  5. 5reflex 4B70.3I 80 · C 75 · S 68 · K 60 · ~$0.022 est.
  6. 6jqv68.6I 79 · C 79 · S 75 · K 47 · ~$0.056 est.
  7. 7decision-machine-168.3I 62 · C 70 · S 93 · K 54 · $0.035
  8. 8decider-35b-a3b67.6I 80 · C 72 · S 81 · K 45 · ~$0.067 est.
  9. 9open-alternative-jev (Qwen3.5-4B)67.0I 64 · C 63 · S 83 · K 60 · ~$0.022 est.
  10. 10system-one-open66.6I 70 · C 57 · S 77 · K 65 · ~$0.015 est.
  11. 11OpenJev (razorback16)66.4I 79 · C 65 · S 83 · K 45 · ~$0.066 est.
  12. 12SimpleJev Qwen3.8-27B66.3I 85 · C 81 · S 71 · K 39 · ~$0.104 est.
  13. 13ZeroEntropy zerank-266.0I 63 · C 76 · S 79 · K 50 · $0.047
  14. 14GPT-5.6 Luna (low)65.9I 95 · C 90 · S 78 · K 28 · $0.242
  15. 15openjev-sglang65.3I 83 · C 77 · S 77 · K 36 · ~$0.131 est.
  16. 16Qwen3-Reranker-4B63.8I 64 · C 67 · S 79 · K 49 · $0.050
  17. 17reflex-27b63.3I 86 · C 86 · S 67 · K 32 · ~$0.181 est.
  18. 18LitJev62.7I 82 · C 84 · S 67 · K 34 · ~$0.163 est.
  19. 19kev 0.6B62.5I 52 · C 51 · S 76 · K 76 · ~$0.0063 est.
  20. 20SimpleJev Qwen3.6-35B-A3B62.5I 80 · C 67 · S 75 · K 38 · ~$0.116 est.
  21. 21djev62.4I 81 · C 93 · S 75 · K 27 · ~$0.274 est.
  22. 22jev-local61.8I 71 · C 69 · S 69 · K 43 · ~$0.077 est.
  23. 23decider-2b61.7I 61 · C 47 · S 83 · K 61 · ~$0.020 est.
  24. 24Bespoke Nimble 9B60.5I 78 · C 65 · S 79 · K 33 · ~$0.166 est.
  25. 25Gemini 3.1 Flash-Lite60.1I 86 · C 68 · S 82 · K 27 · $0.264
  26. 26OpenJev60.0I 88 · C 70 · S 76 · K 28 · ~$0.255 est.
  27. 27kev 4B59.7I 65 · C 42 · S 76 · K 62 · ~$0.019 est.
  28. 28DeepSeek V4.1 Flash57.5I 94 · C 97 · S 72 · K 17 · $0.594
  29. 29kev 8B56.4I 69 · C 44 · S 75 · K 44 · ~$0.073 est.
  30. 30Open-Jev 9B55.0I 71 · C 63 · S 72 · K 28 · ~$0.249 est.
  31. 31system-one54.8I 70 · C 37 · S 84 · K 41 · ~$0.089 est.
  32. 32jeff54.4I 47 · C 65 · S 63 · K 77 · ~$0.0060 est.
  33. 33Laya54.4I 46 · C 62 · S 71 · K 86 · ~$0.0029 est.
  34. 34Open-Jev 2B51.3I 61 · C 55 · S 73 · K 28 · ~$0.249 est.
  35. 35OpenDecision40.6I 41 · C 56 · S 80 · K 75 · ~$0.0066 est.
  36. 36openJev Verdict 1.438.9I 39 · C 74 · S 78 · K 82 · ~$0.0039 est.
  37. 37openJev Verdict38.1I 40 · C 51 · S 77 · K 83 · ~$0.0037 est.
  38. 38kev 0.5B33.2I 38 · C 47 · S 77 · K 76 · ~$0.0063 est.
  39. 39GLiNER2 large29.6I 40 · C 24 · S 62 · K 73 · ~$0.0077 est.
  40. 40smalljev semantic-v927.4I 35 · C 59 · S 80 · K 58 · ~$0.025 est.
  41. 41GLiNER224.0I 36 · C 24 · S 72 · K 83 · ~$0.0037 est.
  42. 42open-jev-deberta-v3-large23.1I 32 · C 66 · S 66 · K 74 · ~$0.0073 est.
  43. 43GLiNER2.5 multi16.6I 28 · C 56 · S 68 · K 82 · ~$0.0039 est.
  44. 44GLiNER2.5 small13.8I 26 · C 47 · S 78 · K 82 · ~$0.0039 est.
  45. 45Mixedbread mxbai-rerank-base-v20.8I 7 · C 83 · S 88 · K 68 · $0.012
  46. 46BAAI bge-reranker-v2-m30.7I 6 · C 84 · S 90 · K 73 · $0.0077
  47. 47Alibaba GTE Reranker ModernBERT-base0.3I 5 · C 77 · S 91 · K 70 · $0.010
  48. 48Certo v10.0I 0 · C 82 · S 94 · K 100 · ~$0.0010 est.
  49. classifier.dev (fast tier) (honorable mention)83.6I 85 · C 78 · S 88 · K 84 · ~$0.0033 est.
  50. Qwen3.8 27B (partial run)24.8I 67 · C 92 · S 61 · K 0 · ~$2.669 est.
  51. Needle 3, options as tools (partial run)1.1I 14 · C · S 53 · K 65 · ~$0.014 est.
  52. Needle 3 (partial run)0.1I 5 · C · S 60 · K 59 · ~$0.024 est.

Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)

  • Jev (TypeSafe, closed)
  • Jev rebuild (open, or open source planned)
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier (not a Jev rebuild)
  • Closed decision model (API only, not Jev)
  • Shown, not ranked — honorable mention (runs another entrant's model) · partial run
Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann. = estimated/announced cost; † = see note.
Legend and notes
  • ~ est. = no measured bill; priced like a large inference provider (how costs are estimated).
  • ann. = the provider’s announced price, not yet charged.
  • Names link to each project.
  • A label-only system has no calibration (–, counted as 0).
  • djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • open-alternative-jev (Qwen3.5-4B): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.

Weighting: Intelligence : Calibration : Speed : Cost

Official default

Custom

The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.

Explore by task difficulty

All tasks is the published default. Choose a scope to see how the ranking changes by difficulty. Hard only uses all 220 hard-tier decisions and their measured Intelligence, Calibration, Speed and Cost.

Tier mapping: Easy = easy; Medium = standard. Easy scopes change Intelligence only. Hard only measures all four axes on the same hard-tier subset; systems without a hard-tier run are shown as partial and are not ranked.

What the run says (JevBench Score)

  • Jev 1.13.0 (TypeSafe AI) leads with 74.4: Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
  • classifier.dev scores 83.6 — higher than anything in the ranking — but is not ranked: it runs Jev (TypeSafe), so ranking it would put the same model in the list twice, once at the model's own price and once at the service's. It keeps every number it earned under Honorable mentions — services built on another entrant's model.
  • Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 73.11.3 points behind: more speed and a lower (estimated) price, less intelligence and calibration.
  • GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (95.3) but places #14: its cost score is 28.5 ($0.242 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.
  • Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not answer every tier — each for the reason in its † note; they are shown below the ranking as partial runs, without a rank.

Axes, tiers, latency and cost

Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.

Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.💲 $ per 1,000 decisions, not per 1,000 tokens — one decision ≈ 950 input tokens.
Rank#SystemEndpoint
1by TypeSafe AIJev 1.13.074.485.782.783.352.0$0.040100.0%99.0%94.5%74.1%0.65 s rawp95 0.72 s rawproduction API
2by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ73.179.072.683.759.5~$0.022 est.100.0%97.9%95.2%59.5%0.20 s raw0.55 s adjustedp95 0.32 s raw → 0.78 sour RunPod GPU
3by Maisa (David Villalón)djevMaisa, diffusion-gemma73.082.765.491.457.6$0.026 announced100.0%97.9%93.2%69.5%0.24 s rawp95 0.31 s rawproduction API
4by Eldan RingWinnow-12B Q871.282.072.082.352.9~$0.037 est.100.0%96.9%91.1%70.9%0.23 s raw0.60 s adjustedp95 0.41 s raw → 0.98 sour RunPod GPU
5by kshetrajna12reflex 4B70.380.175.268.059.7~$0.022 est.100.0%94.8%97.3%63.2%1.80 s raw3.75 s adjustedp95 2.05 s raw → 4.26 sour RunPod GPU
6by hjmurmur (Octalab)jqvQwen3-32B zero-shot68.679.379.074.647.5~$0.056 est.100.0%95.8%92.5%64.5%0.75 s raw1.64 s adjustedp95 0.97 s raw → 2.10 sour RunPod GPU
7by milliseconds.ai (Baptiste Laget)decision-machine-1milliseconds.ai68.362.170.492.953.7$0.035100.0%76.0%89.7%46.8%0.17 s rawp95 0.30 s rawproduction API
8by Mapikadecider-35b-a3b67.679.671.580.845.3~$0.067 est.100.0%96.9%91.1%65.5%0.29 s raw0.73 s adjustedp95 0.49 s raw → 1.14 sour RunPod GPU
9by IkerMoelopen-alternative-jevQwen3.5-4B, IkerMoel67.064.063.283.559.6~$0.022 est.100.0%84.4%74.7%56.8%0.21 s raw0.56 s adjustedp95 0.32 s raw → 0.80 sour RunPod GPU
10by mithalounisystem-one-openGemma 4 E2B LoRA on an L466.669.556.777.064.8~$0.015 est.100.0%93.8%87.7%49.1%0.65 s raw1.30 s adjustedp95 0.77 s raw → 1.54 sauthor's demo server
11by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback1666.479.264.883.245.5~$0.066 est.100.0%95.8%91.1%65.5%0.24 s raw0.63 s adjustedp95 0.31 s raw → 0.76 sour RunPod GPU
12by Featherless AISimpleJev Qwen3.8-27B66.384.781.171.239.5~$0.104 est.100.0%96.9%93.2%75.0%1.01 s raw2.03 s adjustedp95 1.88 s raw → 3.76 sauthor's demo server
13by ZeroEntropyZeroEntropy zerank-266.063.076.579.049.8$0.047100.0%79.2%88.4%47.3%0.13 s raw0.40 s adjustedp95 1.50 s raw → 3.15 sour RunPod GPU
14by OpenAIGPT-5.6 Lunalow reasoning effort65.995.389.877.528.5$0.242100.0%97.9%96.6%94.5%0.97 s rawp95 1.82 s rawproduction API
15by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang65.383.477.477.136.5~$0.131 est.100.0%95.8%95.2%71.4%0.68 s raw1.36 s adjustedp95 0.73 s raw → 1.45 sauthor's demo server
16by QwenQwen3-Reranker-4B63.864.067.078.749.2$0.050100.0%79.2%87.7%50.0%0.13 s raw0.41 s adjustedp95 1.56 s raw → 3.27 sour RunPod GPU
17by kshetrajna12reflex-27bQwen3.8-27B63.385.886.267.532.3~$0.181 est.100.0%95.8%95.9%75.9%1.89 s raw3.93 s adjustedp95 2.21 s raw → 4.57 sour RunPod GPU
18by Zhengxu YuLitJevQwen3.8-27B62.782.483.566.733.6~$0.163 est.100.0%97.9%88.4%73.2%2.03 s raw4.20 s adjustedp95 2.46 s raw → 5.06 sour RunPod GPU
19by Jared Palmerkev 0.6Bresearch preview62.551.951.175.676.1~$0.0063 est.100.0%81.3%66.4%40.0%0.59 s raw1.33 s adjustedp95 0.97 s raw → 2.09 sour RunPod GPU
20by Featherless AISimpleJev Qwen3.6-35B-A3B62.579.567.175.038.1~$0.116 est.100.0%93.8%93.2%66.4%0.85 s raw1.70 s adjustedp95 0.93 s raw → 1.86 sauthor's demo server
21by David Villalon / Maisadjevthinking62.480.892.775.226.9~$0.274 est.95.8%99.0%80.1%77.7%0.43 s raw1.00 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
22by us (GitHub)jev-localQwen3.5-9B61.870.868.769.243.3~$0.077 est.100.0%84.4%89.0%59.1%1.05 s raw2.24 s adjustedp95 2.62 s raw → 5.38 sour RunPod GPU
23by Mapikadecider-2b61.761.246.683.261.0~$0.020 est.100.0%85.4%77.4%47.3%0.26 s raw0.67 s adjustedp95 0.28 s raw → 0.72 sour RunPod GPU
24by Bespoke LabsBespoke Nimble 9B60.577.965.378.733.4~$0.166 est.100.0%94.8%89.0%65.5%0.39 s raw0.93 s adjustedp95 0.65 s raw → 1.46 sour RunPod GPU
25by GoogleGemini 3.1 Flash-Lite60.185.668.181.827.4$0.264100.0%99.0%93.2%75.0%0.76 s rawp95 0.88 s rawproduction API
26by razorback16OpenJevthinking, BF1660.088.069.676.127.8~$0.255 est.100.0%100.0%94.5%78.2%0.46 s raw1.08 s adjustedp95 1.08 s raw → 2.31 sour RunPod GPU
27by Jared Palmerkev 4Bresearch preview59.764.842.075.761.8~$0.019 est.100.0%91.7%85.6%42.3%0.55 s raw1.25 s adjustedp95 0.99 s raw → 2.13 sour RunPod GPU
28by DeepSeekDeepSeek V4.1 Flashthinking default57.594.396.771.616.8$0.59498.6%99.0%93.2%95.0%1.42 s rawp95 4.89 s rawproduction API
29by Jared Palmerkev 8Bresearch preview56.469.444.274.944.0~$0.073 est.100.0%92.7%90.4%47.3%0.59 s raw1.33 s adjustedp95 1.15 s raw → 2.45 sour RunPod GPU
30by Zefan Cai (@Zefan_Cai)Open-Jev 9BZefan Cai55.071.263.372.028.1~$0.249 est.100.0%90.6%81.5%60.9%0.75 s raw1.66 s adjustedp95 1.81 s raw → 3.77 sour RunPod GPU
31by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke54.870.336.884.441.5~$0.089 est.100.0%90.6%91.8%50.0%0.17 s raw0.48 s adjustedp95 0.30 s raw → 0.76 sour RunPod GPU
32by Logan MarkewichjeffLogan Markewich, GLiFormer 400M54.446.964.663.576.6~$0.0060 est.100.0%76.0%61.6%37.7%0.94 s raw2.03 s adjustedp95 10.97 s raw → 22.09 sour CPU
33by Convai InnovationsLayaConvai Innovations, ModernBERT-large 421M54.445.862.571.186.2~$0.0029 est.94.4%72.9%69.2%34.1%0.79 s raw1.72 s adjustedp95 2.20 s raw → 4.54 sour CPU
34by Zefan Cai (@Zefan_Cai)Open-Jev 2BZefan Cai51.361.055.173.528.1~$0.249 est.100.0%79.2%88.4%42.7%0.66 s raw1.48 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
35by Deepan WadhwaOpenDecisionModernBERT-large zero-shot40.640.856.179.975.3~$0.0066 est.87.5%62.5%71.2%33.2%0.34 s raw0.83 s adjustedp95 0.54 s raw → 1.24 sour RunPod GPU
36by Hemant (heman10x)openJev Verdict 1.438.938.674.178.182.4~$0.0039 est.86.1%67.7%56.2%37.7%0.31 s raw0.78 s adjustedp95 0.92 s raw → 2.00 sour CPU
37by Hemant (heman10x)openJev Verdictheman10x, ModernBERT-base 151M38.139.851.376.783.1~$0.0037 est.86.1%65.6%61.0%38.2%0.28 s raw0.71 s adjustedp95 1.45 s raw → 3.04 sour CPU
38by Jared Palmerkev 0.5B33.238.247.477.076.1~$0.0063 est.95.8%52.1%71.2%30.9%0.43 s raw1.01 s adjustedp95 0.92 s raw → 1.99 sour RunPod GPU
39by FastinoGLiNER2 large29.640.124.361.773.3~$0.0077 est.98.6%62.5%61.0%36.4%1.10 s raw2.34 s adjustedp95 14.49 s raw → 29.13 sour CPU
40by Aditya (isHeSatoshi)smalljev semantic-v927.435.158.979.857.9~$0.025 est.97.2%68.8%40.4%38.2%0.41 s raw0.98 s adjustedp95 0.46 s raw → 1.07 sour RunPod GPU
41by FastinoGLiNER2Fastino, gliner2.5-base24.035.623.771.883.1~$0.0037 est.97.2%66.7%45.9%36.4%0.31 s raw0.78 s adjustedp95 4.15 s raw → 8.46 sour CPU
42by Kotoba Labsopen-jev-deberta-v3-largelocal CPU23.131.966.466.074.0~$0.0073 est.100.0%49.0%53.4%36.4%1.77 s raw3.69 s adjustedp95 3.35 s raw → 6.85 sour CPU
43by FastinoGLiNER2.5 multiFastino, 287M16.627.756.167.882.4~$0.0039 est.90.3%51.0%43.8%37.7%0.43 s raw1.01 s adjustedp95 8.18 s raw → 16.50 sour CPU
44by FastinoGLiNER2.5 smallFastino, 74M13.825.647.277.882.4~$0.0039 est.83.3%47.9%50.0%33.2%0.11 s raw0.38 s adjustedp95 2.10 s raw → 4.35 sour CPU
45by MixedbreadMixedbread mxbai-rerank-base-v20.86.783.187.567.9$0.01244.4%33.3%26.7%40.0%0.07 s raw0.29 s adjustedp95 0.23 s raw → 0.62 sour RunPod GPU
46by BAAIBAAI bge-reranker-v2-m30.76.383.889.573.4$0.007743.1%36.5%8.9%36.8%0.03 s raw0.22 s adjustedp95 0.18 s raw → 0.51 sour RunPod GPU
47by Alibaba-NLPAlibaba GTE Reranker ModernBERT-base0.34.676.890.669.6$0.01033.3%39.6%30.1%33.6%0.05 s raw0.25 s adjustedp95 0.10 s raw → 0.35 sour RunPod GPU
48by AltSlate LabsCerto v10.00.082.094.0100.0~$0.0010 est.27.8%30.2%21.9%31.8%0.02 s raw0.19 s adjustedp95 0.03 s raw → 0.21 sour RunPod GPU
Honorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.
by mrmps (@michael_chomsky)classifier.devfast tierhonorable mention · not ranked83.685.177.987.684.3~$0.0033 est.100.0%99.0%97.3%70.5%0.39 s rawp95 0.45 s rawproduction API
Partial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions.
by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked24.867.492.161.30.0~$2.669 est.98.6%99.0%95.3%21.4%5.75 s rawp95 12.97 s rawproduction API
by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked1.113.5none (label only)52.865.3~$0.014 est.66.7%31.3%34.2%3.78 s raw7.71 s adjustedp95 33.64 s raw → 67.42 sour CPU
by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked0.14.6none (label only)59.958.7~$0.024 est.47.2%16.7%31.5%7.7%1.69 s raw3.52 s adjustedp95 14.36 s raw → 28.88 sour CPU
† Notes on 39 marked systems — how each was run
  • djev: The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • classifier.dev: Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.

Honorable mentions — services built on another entrant's model

A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.

classifier.devno rank

83.6 JevBench Score · Jev 1.13.0 (#1) scores 74.4

Runs on Jev (TypeSafe).

  • Intelligence85.1
  • Calibration77.9
  • Speed87.6
  • Cost84.3
  • $ per 1,000 decisions~$0.0033 est.

Runs on Jev (TypeSafe) — listed, not ranked. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.

Why it is not ranked, what its price assumes, and what we found

classifier.dev is not its own model. Its own pages say so: "The fast tier is Jev, TypeSafe's decision model" (https://classifier.dev/benchmark, read 2026-09-20), and the API answers with "model": "jev-1.13.0" — the same model version this benchmark measures directly as Jev 1.13.0. What it adds is a price and, on its smart tier, an orchestration layer: "The smart tier is Jev plus a reasoning model re-asking only the answers Jev put under 0.7 confidence" — escalation on low confidence (a model cascade), not best-of-N, not self-consistency and not a committee. Its published escalation model is gemini-3.8-flash. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.

Only the fast tier was measured. The smart tier's escalation was never run, so nothing here scores it.

Price. $0.0033 per 1,000 decisions is an estimate from the published flat-rate plan at full use: classifier.dev Pro is $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-20), and one classification is one decision. Lower use costs more per decision — at a tenth of that allowance it is $0.033 per 1,000 — and the free tier (20,000 fast classifications a day), which is what our run used, costs nothing. Their pages do not say how the flat rate is funded, so we do not know their cost basis; the only figure they publish is what the model costs a caller: "The model behind the fast tier costs about $0.005 per thousand classifications and needs a TypeSafe key" (https://classifier.dev/pricing) — for their short single-sentence inputs, not for JevBench's whole questions.

Not a pass-through. On our set the fast tier scored 97.3 % on the judge tier against Jev's 94.5 %, and 70.5 % against 74.1 % on the hard tier. classifier.dev's own explanation for differences of this kind is batching ("The fast tier is Jev, packed a thousand to a request"); on their own two test sets they measured the same difference as noise.

A legitimate, well-documented product: free without an account, open source (https://github.com/mrmps/classifier-dev), by Michael Ryaboy (@michael_chomsky). Read 2026-09-20: classifier.dev · classifier.dev/benchmark · classifier.dev/pricing · classifier.dev/about

Which public tasks did each system get right?

This view shows public task outcomes only: 231 of 231 public tasks in the selected scope. Held-out and imported task text is not shipped.

Show 231 public task outcomes across 52 systems
TaskJev 1.13.0SemIfdjevWinnow-12B Q8reflex 4Bjqvdecision-machine-1decider-35b-a3bopen-alternative-jevsystem-one-openOpenJevSimpleJev Qwen3.8-27BZeroEntropy zerank-2GPT-5.6 Lunaopenjev-sglangQwen3-Reranker-4Breflex-27bLitJevkev 0.6BSimpleJev Qwen3.6-35B-A3Bdjevjev-localdecider-2bBespoke Nimble 9BGemini 3.1 Flash-LiteOpenJevkev 4BDeepSeek V4.1 Flashkev 8BOpen-Jev 9Bsystem-onejeffLayaOpen-Jev 2BOpenDecisionopenJev Verdict 1.4openJev Verdictkev 0.5BGLiNER2 largesmalljev semantic-v9GLiNER2open-jev-deberta-v3-largeGLiNER2.5 multiGLiNER2.5 smallMixedbread mxbai-rerank-base-v2BAAI bge-reranker-v2-m3Alibaba GTE Reranker ModernBERT-baseCerto v1classifier.devQwen3.8 27BNeedle 3, options as toolsNeedle 3
Easy · 48 of 72 decisions publicEasy · 48 of 72 public48/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4842/4842/4841/4846/4848/4847/4847/4848/4844/4841/4822/4823/4817/4812/4848/4848/4832/4823/48
easy-intent-00choice××××××
easy-intent-01choice××××
easy-intent-02choice××××××
easy-intent-03choice×××
easy-intent-04choice××××
easy-intent-05choice×××××
easy-intent-06choice×××××
easy-intent-07choice×××
easy-intent-08choice××××××
easy-intent-09choice×××××
easy-intent-10choice×××
easy-intent-11choice××
easy-fact-00noul××
easy-fact-01noul×××××××××
easy-fact-02noul×××
easy-fact-03noul××××××××
easy-fact-04noul××××
easy-fact-05noul××××××××
easy-fact-06noul×××
easy-fact-07noul×××××××
easy-fact-08noul××
easy-fact-09noul××××××××
easy-fact-10noul××××
easy-fact-11noul×××××××
easy-extraction-00choice×××××
easy-extraction-01choice××××
easy-extraction-02choice×××
easy-extraction-03choice×××
easy-extraction-04choice×××××
easy-extraction-05choice×
easy-extraction-06choice××××
easy-extraction-07choice××
easy-extraction-08choice××
easy-extraction-09choice××
easy-extraction-10choice××××
easy-extraction-11choice××
easy-tool_selection-00choice××
easy-tool_selection-01choice××××
easy-tool_selection-02choice×××
easy-tool_selection-03choice!××××
easy-tool_selection-04choice××××
easy-tool_selection-05choice××××
easy-tool_selection-06choice××××
easy-tool_selection-07choice!×××
easy-tool_selection-08choice××××
easy-tool_selection-09choice××××
easy-tool_selection-10choice××
easy-tool_selection-11choice×××
Medium (standard) · 72 of 96 decisions publicMedium (standard) · 72 of 96 public71/7271/7271/7269/7268/7269/7254/7270/7260/7267/7270/7270/7257/7270/7268/7254/7269/7271/7258/7267/7271/7260/7261/7267/7271/7272/7264/7271/7267/7265/7264/7254/7250/7255/7243/7250/7245/7235/7242/7249/7246/7231/7232/7230/7224/7226/7226/7224/7271/7271/7219/7212/72
original-policy-01-0noul××××××
original-policy-01-1noul!××××××××
original-policy-02-0noul××××××××××
original-policy-02-1noul××××××××××××××××××××
original-policy-03-0noul×××××××××××××××××
original-policy-03-1noul×××××××××××××××××××
original-policy-04-0noul×××××××××××
original-policy-04-1noul×××××××××
original-policy-05-0noul×××××××××
original-policy-05-1noul××××××
original-policy-06-0noul××××××××××
original-policy-06-1noul××××××××××××××××
original-intent-01-0choice××××××××××
original-intent-01-1choice×××××××××××
original-intent-02-0choice×××××××××××××××××××××××
original-intent-02-1choice×××××××××××
original-intent-03-0choice×××××××××××××××
original-intent-03-1choice××××
original-intent-04-0choice××××
original-intent-04-1choice××××××××××
original-intent-05-0choice×××××××××××××××××
original-intent-05-1choice×××××××××××××××××
original-intent-06-0choice××××××××××××
original-intent-06-1choice××××××××××××
original-ordinal-01-0score××××××××××
original-ordinal-01-1score×××××××
original-ordinal-02-0score××××××××
original-ordinal-02-1score××××××××
original-ordinal-03-0score×××××××
original-ordinal-03-1score×××××××××
original-ordinal-04-0score××××××
original-ordinal-04-1score××××××
original-ordinal-05-0score×××!×××××××××××
original-ordinal-05-1score××××××××××××
original-ordinal-06-0score×××××
original-ordinal-06-1score××××
original-extraction-01-0choice×××××××××××××××××
original-extraction-01-1choice××××××××××××××
original-extraction-02-0choice×××××××
original-extraction-02-1choice×××××××
original-extraction-03-0choice××××××××××××××
original-extraction-03-1choice××××××××××
original-extraction-04-0choice×××××××××××
original-extraction-04-1choice×××××××××
original-extraction-05-0choice×××××××
original-extraction-05-1choice×××××××
original-extraction-06-0choice××××××
original-extraction-06-1choice××××××××××
original-adequacy-01-0noul××××××××××
original-adequacy-01-1noul×××××××
original-adequacy-02-0noul×××××××××××××××××
original-adequacy-02-1noul××××××
original-adequacy-03-0noul××××××××××××××××××××××××××××××××
original-adequacy-03-1noul××××××××××××××××××××××××××××××××
original-adequacy-04-0noul××××××××××
original-adequacy-04-1noul××××××××××××
original-adequacy-05-0noul×××××××××××××××××××××××××××
original-adequacy-05-1noul×××××××××××××××××××××××××
original-adequacy-06-0noul×××××××
original-adequacy-06-1noul×××××××××××
original-routing-01-0choice××××××
original-routing-01-1choice×××××××××××××××
original-routing-02-0choice××××××××
original-routing-02-1choice×××××××××
original-routing-03-0choice×××××
original-routing-03-1choice×××××
original-routing-04-0choice×××××××××××××××
original-routing-04-1choice××××××××××××××××××××××
original-routing-05-0choice×××××××××××××××××
original-routing-05-1choice××××××××××××××××××××××××
original-routing-06-0choice×××××××××××!×
original-routing-06-1choice×××××××××××××
Hard · 111 of 220 decisions publicHard · 111 of 220 public81/11168/11175/11181/11167/11168/11154/11174/11163/11154/11171/11182/11157/111107/11181/11155/11184/11180/11148/11173/11185/11165/11155/11169/11182/11185/11141/111107/11150/11166/11154/11143/11139/11146/11138/11141/11142/11133/11141/11144/11141/11142/11137/11135/11140/11142/11135/11137/11178/11147/510/017/44
hard-opus-a-long_policy-01choice×!×××××××××××!××××××·×
hard-opus-a-long_policy-04choice×××××××××××××××××××!×××××××××××××××××××××××××××·
hard-opus-a-long_policy-08choice××××××××××××××××××××××·×
hard-opus-a-long_policy-09choice××××××××××××××××××××××××·×
hard-opus-a-long_policy-11noul××××××××××××××××××·×
hard-opus-a-long_policy-13noul×××××××××××××××××××××××××××××!·
hard-opus-a-long_policy-17choice××××××××!×××××××××××××××××××·×
hard-opus-a-long_policy-19noul××××××××××××××××××××××××××××××·×
hard-opus-a-probability-03noul××××××××××××××·
hard-opus-a-probability-04choice××××××××××××××××××××××××××××××××××××××××××·×
hard-opus-a-probability-07choice××××××××××××××××××××××××××××××·×
hard-opus-a-probability-08noul××××××××××××××××××××·×
hard-opus-a-temporal_numeric-03noul×××××××××××××××××××××××××××××××××××·
hard-opus-a-temporal_numeric-06noul××××××××××××××××××××·×
hard-opus-a-temporal_numeric-07choice×××××××××××××××××××!××××××××××××××××·
hard-opus-a-temporal_numeric-09noul×××××××××××××××××××××××××·×
hard-opus-a-temporal_numeric-12score××××××××××××××××!××××××××××××××××××××××××·
hard-opus-b-ambiguous-02choice××××××××××××××××·×
hard-opus-b-ambiguous-03choice×××××××××××××!×××××××××××××·
hard-opus-b-ambiguous-07choice×××××××××××××!××××××××××××××××××·
hard-opus-b-ambiguous-09choice××××××××××××××××××××××××××·
hard-opus-b-ambiguous-10choice×××××××××××××××××××××·×
hard-opus-b-ambiguous-11choice×××××××××××××××××·×
hard-opus-b-ambiguous-13choice××!×××××××××××××××××·×
hard-opus-b-multi_hop-03choice××××××××××!×××××××××××××××××××××·×
hard-opus-b-multi_hop-04choice××××××××××××××××××××××××××××××××××××××××××·×
hard-opus-b-multi_hop-05score×××××××××××××××××××××××××××××××××××××××××·×
hard-opus-b-multi_hop-07choice××××××××××!×××××××××××××××××××××·×
hard-opus-b-multi_hop-08noul×××××××××××××××××·×
hard-opus-b-probability-01choice×××××××××××××××××××××·×
hard-opus-b-probability-02choice××××××××××××××××××××××××××××××××××××××××××·×
hard-opus-b-probability-03noul××××××××××××××·
hard-opus-b-probability-04choice××××××××××××××·
hard-opus-b-probability-06choice××××××××××××××!××××××××××××·
hard-opus-b-tradeoff-01choice××××××××××××××××××××××××××××××××××××××·×
hard-opus-b-tradeoff-03choice××××××××××××××××××××·×
hard-opus-b-tradeoff-06noul××××××××××××××××××××·
hard-opus-b-tradeoff-07choice×××××××××××××××××××××××××·×
hard-opus-b-tradeoff-08noul××××××××××××××××××××××××××××××·
hard-opus-b-tradeoff-12choice××!×××××××××·×
hard-opus-c-long_policy-02choice××!××××××××××××·×
hard-opus-c-long_policy-03choice×××××××××××××××××××××××××××××·
hard-opus-c-long_policy-04noul××××××××××·
hard-opus-c-long_policy-05score×××××××××××××××××××××××××××××××××××××××××!·
hard-opus-c-long_policy-08score××××××××××××××××××××××××××××××××××××××××××··
hard-opus-c-long_policy-10choice×××××××××××··
hard-opus-c-long_policy-11noul×××××××××××××××××××××××××××××××××××××!··
hard-opus-c-probability-03choice×××××××××××××××××··
hard-opus-c-temporal_numeric-02noul×××××××××××××××××××××××××××××··
hard-opus-c-temporal_numeric-03choice×××××××××××××××!××××××××××××××××××××××××··
hard-opus-c-temporal_numeric-04choice×××××××××××××××××××!××××××××××××××××××××××··
hard-opus-c-temporal_numeric-06choice×××××××××××××××××××××××××××××××××××××××××××···
hard-opus-c-temporal_numeric-08choice×××××××××××××××××××××××××××××××···
hard-opus-c-temporal_numeric-12choice×××××××××××××××××××××××××···
hard-sol-a-adversarial-01choice×××××!××××××××××××···
hard-sol-a-adversarial-06noul×××···
hard-sol-a-adversarial-07choice×××××××××××···
hard-sol-a-adversarial-08noul××××××××××···
hard-sol-a-adversarial-09score×××××××××···
hard-sol-a-adversarial-11choice×××××××××××××××××···
hard-sol-a-multi_hop-01choice××××××××···
hard-sol-a-multi_hop-02choice×××××××××!××××××××××××××××···
hard-sol-a-multi_hop-05choice××××××××××××××××××××××···
hard-sol-a-multi_hop-07choice××××××××××××××××××××××···
hard-sol-a-multi_hop-08choice×××××××××××××××××××××××××××××××××××···
hard-sol-a-multi_hop-09choice××××××××××××××···
hard-sol-a-multi_hop-10choice××××××××××!×××××××××××××××××···
hard-sol-a-multi_hop-12choice×××××××××××××××××××××××××···
hard-sol-a-trap-02noul×××××××××···
hard-sol-a-trap-04noul×××××···
hard-sol-a-trap-05score××××××××××××××××××××···
hard-sol-a-trap-06choice!×××××××××××···
hard-sol-a-trap-08choice×××××××××××××××××××···
hard-sol-a-trap-10choice×××××××××××××···
hard-sol-a-trap-13noul×××××××××××···
hard-sol-a-trap-15noul××××××××××···
hard-sol-b-judge_hard-01noul××××···
hard-sol-b-judge_hard-02noul××××××××××××××××××××!×××××××××××××××××××××××···
hard-sol-b-judge_hard-05noul×××××××···
hard-sol-b-judge_hard-08noul××××××××××××××××!×××××××××××××××××××××···
hard-sol-b-judge_hard-10noul××××××××××××××××××××××××××××××···
hard-sol-b-judge_hard-14noul×××××××××××××××××···
hard-sol-b-judge_hard-15noul×××···
hard-sol-b-judge_hard-18noul××××××××××××××××××××××××···
hard-sol-b-long_policy-01choice×××××!××××××××××××···
hard-sol-b-long_policy-02choice××××××××××××××××××××××××···
hard-sol-b-long_policy-05choice××!××××××××××××××···
hard-sol-b-long_policy-06choice×××××××××××××××!×××××××××××××××···
hard-sol-b-routing_hard-01choice×××××××···
hard-sol-b-routing_hard-02choice×××××××××××···
hard-sol-b-routing_hard-03choice×××××××××××××××···
hard-sol-b-routing_hard-07choice××××××××××××××···
hard-sol-b-routing_hard-09choice×××××××××××××···
hard-sol-b-temporal_numeric-01choice×××××××××××××××××××××××××××××××××××××××××××××···
hard-sol-b-temporal_numeric-02choice×××××××××××××××××××××××××××···
hard-sol-b-temporal_numeric-03choice××××××××××××××!×××××××××××××××××···
hard-sol-b-temporal_numeric-04choice×××××××××××××××××××××××××××××···
hard-sol-c-judge_hard-03noul×××××××××××××××××××××××···
hard-sol-c-judge_hard-05noul×××××××××××××××···
hard-sol-c-judge_hard-07noul××××××××××××××××××××××××××···
hard-sol-c-judge_hard-08noul××××××××××××××××××××××××···
hard-sol-c-judge_hard-09noul×××××××···
hard-sol-c-judge_hard-10noul××××···
hard-sol-c-judge_hard-11noul××××××××××××××××···
hard-sol-c-judge_hard-13noul××××××××××××××××××××××××···
hard-sol-c-judge_hard-15noul××××××××···
hard-sol-c-multi_hop-07choice×××××××××××××···
hard-sol-c-multi_hop-09choice××!××××××××××××···
hard-sol-c-multi_hop-11choice×××!×××××××××××××××···
hard-sol-c-multi_hop-12choice×××××××××××××××××××××××××···
hard-sol-c-multi_hop-13choice×××××××××···

✓ correct · × wrong · ! failed (scored wrong) · · not attempted · — no public outcome in the pinned artifact. A group row counts the public tasks it lists; the tier columns of the table above use all decisions of the tier. Every task id carries its topic (hard-opus-a-long_policy-01 is a long_policy task), and the tag after the id is the published question type: choice — pick one of a defined set of options · noul — whether a stated condition holds · score — a degree along a described dimension. Task ids are shown without their tier prefix; the tier is the group row. Task descriptions are intentionally not included; the task id, tier, topic and type are the published public metadata.

How the JevBench Score works

JevBench Score = (Intelligence × Calibration × Speed × Cost)1/4, each axis on 0–100 — the geometric mean. A weak axis pulls the score down hard: a strong axis cannot buy it back. Below 50 Intelligence the score is also multiplied by (Intelligence ÷ 50)², so a system barely better than guessing cannot rank on speed and price.

  • Intelligence — accuracy above chance: per tier, how much of the gap between guessing and all-correct a system closes (0 = guessing, 100 = all correct), weighted: hard 30 %, easy 14 %, standard 28 %, judge 28 % (220 / 72 / 96 / 146 decisions).
  • Calibration — on the hard tier: does “80 % sure” come true 80 % of the time, and does the returned distribution match the exact gold distribution on the probability items.
  • Speed — median and 95th-percentile latency, one request at a time: 0.1 s scores 100, each 10× slower costs 20 points (1 s = 80, 10 s = 60). Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.
  • Cost — dollars per 1,000 decisions, never per 1,000 tokens: $0.001 scores 100, each 10× more expensive costs 30 points ($0.01 = 70, $0.10 = 40, $1 = 10). Models without a tariff are priced at hosted-provider prices, marked “est.” (how). 💲 $ per 1,000 decisions, not $ per 1,000 tokens. One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.
Full scoring rules

JevBench Score.
Geometric mean of Intelligence, Calibration, Speed and Cost, 25 % each. If chance-corrected Intelligence is below 50, multiply by (Intelligence / 50)^2; at or above 50 there is no penalty.

Intelligence.
Per tier: 100 x (accuracy - chance) / (1 - chance), clipped at 0. Chance is 1 / options for each item (1 / levels for score items), then averaged within the tier. Tier weights: hard 30 %, easy 14 %, standard 28 %, judge 28 %. Failed, timed-out or unparseable answers count as wrong.

Calibration.
Hard tier only, systems that return a probability distribution: mean of (a) 100 x (1 - ECE/0.5), ECE = top-label expected calibration error in 10 bins, and (b) probability fidelity = 100 x (1 - mean total-variation distance) between the returned distribution and the exact gold distribution on the 20 probability items. Label-only systems have none; it counts as 0 in the JevBench Score.

Speed.
Mean of score(p50) and score(p95) of the serial 242-decision standard+judge run; score(s) = 100 - 20 log10(s / 0.1 s), clipped to 0..100 (0.1 s = 100, 1 s = 80, 10 s = 60). Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo. Production APIs (Jev, djev, classifier.dev, OpenAI, Google, DeepSeek, Chutes) are not adjusted.

Cost.
US dollars per 1,000 DECISIONS — not per 1,000 tokens. One decision is one whole question: its state, its rubric and its options, which is hundreds to thousands of input tokens. Pooled over all 534 v1.2 decisions; score = 100 - 30 log10(usd / 0.001), clipped to 0..100 ($0.001 = 100, $0.01 = 70, $0.10 = 40, $1 = 10). Measured = public tariff x measured tokens. est. = hosted-provider list price of the same weights or size class x tokens (for a flat-rate service, its published plan price at full use). announced = the provider's published price, not yet charged (free preview), x measured tokens.

Ranked.
Ranked: a system's own model, with every tier attempted for >= 95 % of its decisions. Partial runs are shown below the ranking, marked, without a rank. A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number.

Presets.
Other views reweight the same four axes and combine them the same way (geometric mean). They are not the JevBench Score.

How the ranking moves with other weights

Rank and score under the JevBench Score and the earlier views, all combined as a geometric mean (Intelligence : Calibration : Speed : Cost). Highlighted = a different rank than the JevBench Score. Ranked systems only.

SystemJevBench Score25:25:25:25 · officialBalanced, no calibration33:0:33:33 · not the defaultEmphasis on Accuracy60:0:20:20 · not the defaultEmphasis on Speed20:0:60:20 · not the defaultEmphasis on Cost20:0:20:60 · not the default
Jev 1.13.0#1 74.4#3 71.8 (rank differs from the JevBench Score)#2 77.1 (rank differs from the JevBench Score)#4 76.2 (rank differs from the JevBench Score)#9 63.1 (rank differs from the JevBench Score)
SemIf#2 73.1#2 73.3#3 75.5 (rank differs from the JevBench Score)#2 77.3#4 67.4 (rank differs from the JevBench Score)
djev#3 73.0#1 75.8 (rank differs from the JevBench Score)#1 78.5 (rank differs from the JevBench Score)#1 81.7 (rank differs from the JevBench Score)#3 67.9
Winnow-12B Q8#4 71.2#4 71.0#4 75.2#5 75.3 (rank differs from the JevBench Score)#10 63.1 (rank differs from the JevBench Score)
reflex 4B#5 70.3#6 68.8 (rank differs from the JevBench Score)#5 73.1#17 68.4 (rank differs from the JevBench Score)#5 65.0
jqv#6 68.6#14 65.5 (rank differs from the JevBench Score)#9 70.7 (rank differs from the JevBench Score)#14 69.0 (rank differs from the JevBench Score)#14 57.6 (rank differs from the JevBench Score)
decision-machine-1#7 68.3#9 67.6 (rank differs from the JevBench Score)#22 65.4 (rank differs from the JevBench Score)#3 76.8 (rank differs from the JevBench Score)#11 61.7 (rank differs from the JevBench Score)
decider-35b-a3b#8 67.6#13 66.3 (rank differs from the JevBench Score)#8 71.3#10 71.7 (rank differs from the JevBench Score)#18 56.9 (rank differs from the JevBench Score)
open-alternative-jev#9 67.0#7 68.3 (rank differs from the JevBench Score)#17 66.6 (rank differs from the JevBench Score)#6 74.0 (rank differs from the JevBench Score)#8 64.7 (rank differs from the JevBench Score)
system-one-open#10 66.6#5 70.3 (rank differs from the JevBench Score)#11 70.0 (rank differs from the JevBench Score)#9 72.9 (rank differs from the JevBench Score)#2 68.0 (rank differs from the JevBench Score)
OpenJev#11 66.4#11 66.9#7 71.6 (rank differs from the JevBench Score)#8 73.0 (rank differs from the JevBench Score)#15 57.3 (rank differs from the JevBench Score)
SimpleJev Qwen3.8-27B#12 66.3#18 62.0 (rank differs from the JevBench Score)#10 70.2 (rank differs from the JevBench Score)#24 65.5 (rank differs from the JevBench Score)#22 51.8 (rank differs from the JevBench Score)
ZeroEntropy zerank-2#13 66.0#15 62.8 (rank differs from the JevBench Score)#29 62.9 (rank differs from the JevBench Score)#15 68.8 (rank differs from the JevBench Score)#16 57.2 (rank differs from the JevBench Score)
GPT-5.6 Luna#14 65.9#23 59.5 (rank differs from the JevBench Score)#6 71.8 (rank differs from the JevBench Score)#23 66.1 (rank differs from the JevBench Score)#30 44.3 (rank differs from the JevBench Score)
openjev-sglang#15 65.3#19 61.7 (rank differs from the JevBench Score)#12 69.6 (rank differs from the JevBench Score)#18 67.4 (rank differs from the JevBench Score)#24 50.0 (rank differs from the JevBench Score)
Qwen3-Reranker-4B#16 63.8#16 62.8#27 63.3 (rank differs from the JevBench Score)#16 68.7#17 56.9 (rank differs from the JevBench Score)
reflex-27b#17 63.3#26 57.2 (rank differs from the JevBench Score)#16 67.2 (rank differs from the JevBench Score)#28 61.1 (rank differs from the JevBench Score)#27 45.5 (rank differs from the JevBench Score)
LitJev#18 62.7#28 57.0 (rank differs from the JevBench Score)#19 66.0 (rank differs from the JevBench Score)#29 60.7 (rank differs from the JevBench Score)#26 46.1 (rank differs from the JevBench Score)
kev 0.6B#19 62.5#12 66.8 (rank differs from the JevBench Score)#30 60.4 (rank differs from the JevBench Score)#13 70.2 (rank differs from the JevBench Score)#1 70.4 (rank differs from the JevBench Score)
SimpleJev Qwen3.6-35B-A3B#20 62.5#21 61.0 (rank differs from the JevBench Score)#14 67.8 (rank differs from the JevBench Score)#21 66.3 (rank differs from the JevBench Score)#23 50.6 (rank differs from the JevBench Score)
djev#21 62.4#30 54.6 (rank differs from the JevBench Score)#25 63.9 (rank differs from the JevBench Score)#27 62.1 (rank differs from the JevBench Score)#34 41.1 (rank differs from the JevBench Score)
jev-local#22 61.8#22 59.6#26 63.9 (rank differs from the JevBench Score)#26 63.3 (rank differs from the JevBench Score)#21 52.5 (rank differs from the JevBench Score)
decider-2b#23 61.7#8 67.7 (rank differs from the JevBench Score)#23 65.1#7 73.5 (rank differs from the JevBench Score)#7 64.9 (rank differs from the JevBench Score)
Bespoke Nimble 9B#24 60.5#24 58.9#20 65.9 (rank differs from the JevBench Score)#22 66.2 (rank differs from the JevBench Score)#25 47.0 (rank differs from the JevBench Score)
Gemini 3.1 Flash-Lite#25 60.1#25 57.6#15 67.5 (rank differs from the JevBench Score)#20 66.3 (rank differs from the JevBench Score)#32 42.8 (rank differs from the JevBench Score)
OpenJev#26 60.0#27 57.1 (rank differs from the JevBench Score)#13 67.9 (rank differs from the JevBench Score)#25 64.0 (rank differs from the JevBench Score)#31 42.8 (rank differs from the JevBench Score)
kev 4B#27 59.7#10 67.2 (rank differs from the JevBench Score)#18 66.2 (rank differs from the JevBench Score)#12 70.5 (rank differs from the JevBench Score)#6 65.0 (rank differs from the JevBench Score)
DeepSeek V4.1 Flash#28 57.5#34 48.4 (rank differs from the JevBench Score)#28 63.2#33 56.6 (rank differs from the JevBench Score)#40 31.7 (rank differs from the JevBench Score)
kev 8B#29 56.4#20 61.2 (rank differs from the JevBench Score)#24 64.3 (rank differs from the JevBench Score)#19 66.3 (rank differs from the JevBench Score)#19 53.6 (rank differs from the JevBench Score)
Open-Jev 9B#30 55.0#32 52.4 (rank differs from the JevBench Score)#31 59.3 (rank differs from the JevBench Score)#30 59.5#35 40.9 (rank differs from the JevBench Score)
system-one#31 54.8#17 62.6 (rank differs from the JevBench Score)#21 65.6 (rank differs from the JevBench Score)#11 70.6 (rank differs from the JevBench Score)#20 53.1 (rank differs from the JevBench Score)
jeff#32 54.4#31 53.6 (rank differs from the JevBench Score)#33 48.2 (rank differs from the JevBench Score)#34 54.5 (rank differs from the JevBench Score)#13 58.7 (rank differs from the JevBench Score)
Laya#33 54.4#29 55.0 (rank differs from the JevBench Score)#34 47.7 (rank differs from the JevBench Score)#32 56.8 (rank differs from the JevBench Score)#12 61.4 (rank differs from the JevBench Score)
Open-Jev 2B#34 51.3#33 50.1 (rank differs from the JevBench Score)#32 54.2 (rank differs from the JevBench Score)#31 58.4 (rank differs from the JevBench Score)#37 39.8 (rank differs from the JevBench Score)
OpenDecision#35 40.6#35 41.7#35 35.2#35 46.0#28 44.9 (rank differs from the JevBench Score)
openJev Verdict 1.4#36 38.9#37 37.4 (rank differs from the JevBench Score)#38 30.7 (rank differs from the JevBench Score)#37 40.8 (rank differs from the JevBench Score)#33 41.6 (rank differs from the JevBench Score)
openJev Verdict#37 38.1#36 40.1 (rank differs from the JevBench Score)#36 33.3 (rank differs from the JevBench Score)#36 43.3 (rank differs from the JevBench Score)#29 44.7 (rank differs from the JevBench Score)
kev 0.5B#38 33.2#39 35.4 (rank differs from the JevBench Score)#39 29.4 (rank differs from the JevBench Score)#38 38.9#38 38.7
GLiNER2 large#39 29.6#38 36.5 (rank differs from the JevBench Score)#37 31.8 (rank differs from the JevBench Score)#39 37.8#36 40.5 (rank differs from the JevBench Score)
smalljev semantic-v9#40 27.4#41 26.9 (rank differs from the JevBench Score)#41 22.6 (rank differs from the JevBench Score)#41 31.4 (rank differs from the JevBench Score)#41 27.6 (rank differs from the JevBench Score)
GLiNER2#41 24.0#40 30.3 (rank differs from the JevBench Score)#40 24.6 (rank differs from the JevBench Score)#40 32.6 (rank differs from the JevBench Score)#39 34.6 (rank differs from the JevBench Score)
open-jev-deberta-v3-large#42 23.1#42 21.9#42 17.8#42 23.7#42 24.9
GLiNER2.5 multi#43 16.6#43 16.4#43 12.6#43 18.0#43 19.5
GLiNER2.5 small#44 13.8#44 14.4#44 10.6#44 16.5#44 16.9
Mixedbread mxbai-rerank-base-v2#45 0.8#45 0.6#45 0.3#45 0.9#45 0.8
BAAI bge-reranker-v2-m3#46 0.7#46 0.5#46 0.3#46 0.8#46 0.7
Alibaba GTE Reranker ModernBERT-base#47 0.3#47 0.3#47 0.1#47 0.4#47 0.4
Certo v1#48 0.0#48 0.0#48 0.0#48 0.0#48 0.0

Compare two systems

Pick any two. The first radar shows the four axes of the JevBench Score (0–100, the values in the table above); the second shows accuracy by subject topic, over all tiers. Further out is better on every spoke.

  • A: Jev 1.13.0 Jev · JevBench Score 74.4 (#1)
  • B: SemIf Jev rebuild · JevBench Score 73.1 (#2)

The four score axes

Radar: the four JevBench Score axes, two systemsJev 1.13.0 vs SemIf. Intelligence: 85.7 vs 79.0; Calibration: 82.7 vs 72.6; Speed: 83.3 vs 83.7; Cost: 52.0 vs 59.5.50100Intelligence85.7 · 79.0Calibration82.7 · 72.6Speed83.3 · 83.7Cost52.0 · 59.5
Speed includes the latency adjustment for self-hosted and demo endpoints — an assumption, see Limits. A label-only system has no calibration (counted as 0).
Values as a table
AxisA: Jev 1.13.0B: SemIf
Intelligence85.779.0
Calibration82.772.6
Speed83.383.7
Cost52.059.5
JevBench Score74.473.1

Accuracy by subject topic — not part of the score

Radar: accuracy by subject topic, two systemsAccuracy by subject topic, Jev 1.13.0 vs SemIf. Math & numbers (129 items): 87.6% vs 79.1%; Coding & software (56 items): 83.9% vs 96.4%; Rules, policy & law (67 items): 83.6% vs 64.2%; Finance & commerce (64 items): 73.4% vs 60.9%; Support & operations (119 items): 89.1% vs 87.4%; Everyday language (79 items): 100.0% vs 100.0%; Safety & security (20 items): 100.0% vs 75.0%.50100Math87.6% · 79.1%Coding83.9% · 96.4%Rules & law83.6% · 64.2%Finance73.4% · 60.9%Support & ops89.1% · 87.4%Everydaylanguage100.0% · 100.0%Safety &security100.0% · 75.0%
Share of each topic's decisions answered correctly, all tiers together — compare the two systems within a topic, not topics with each other.
Values and notes
  • Topics mix tiers differently — Everyday language is mostly easy items, Rules & law and Finance mostly hard ones — which is why topics are not compared with each other.
  • Topics: one per item, drafted by a model and checked by hand — method. Held-out items count in the totals; their texts stay private.
Topic (items)A: Jev 1.13.0B: SemIf
Math & numbers (129)a calculation decides the answer: arithmetic, word problems, probability, dates, units87.6% 113 of 12979.1% 102 of 129
Coding & software (56)code, SQL, repositories, developer tools and IT systems83.9% 47 of 5696.4% 54 of 56
Rules, policy & law (67)applying written rules: company policies, contracts, regulations, eligibility83.6% 56 of 6764.2% 43 of 67
Finance & commerce (64)money: payments, refunds, invoices, orders, expenses, insurance payouts73.4% 47 of 6460.9% 39 of 64
Support & operations (119)support tickets, incidents, logistics, scheduling desks and routing work to a team89.1% 106 of 11987.4% 104 of 119
Everyday language (79)short everyday messages: intents, assistant requests, reading a detail out of a text100.0% 79 of 79100.0% 79 of 79
Safety & security (20)untrusted or injected instructions, fraud, moderation, access and security triage100.0% 20 of 2075.0% 15 of 20

Jev alternatives, open source and self-hosting

The table above compares the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.

What are open-source alternatives to Jev?

The highest-ranked open entrants in this run are SemIf (#2, 73.1), djev (#3, 73.0), Winnow-12B Q8 (#4, 71.2), reflex 4B (#5, 70.3). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.

Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?

Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.

jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.

How is JevBench scored?

The official score is the geometric mean of chance-corrected Intelligence, Calibration, Speed and Cost, weighted 25% each. Version v1.3.0 uses 534 decisions (220 hard) and penalises systems below 50 Intelligence. Open Method and tiers for the exact rules, or inspect the MIT-licensed harness and public tasks.

How do I submit my model?

Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version. For private data, see the custom evaluation options.

Held-out hard-tier detail

With about 110 items on each side, ordinary noise is roughly ±9 percentage points. Read a system's public-minus-held-out gap against the field mean (-0.7 points across 49 complete systems): only an outlier against that field is meaningful. “Not public” does not mean “not seen”, because held-out items were sent to hosted APIs.

SystemHard publicHard held-outPublic − held-out gap (95% interval)Field mean gap
SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)61.3% (68/111)57.8% (63/109)+3.5 points [-9.5, +16.4]-0.7 points
OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)64.0% (71/111)67.0% (73/109)-3.0 points [-15.6, +9.6]-0.7 points
system-one (Qwen3-8B, Sean Goedecke)48.6% (54/111)51.4% (56/109)-2.7 points [-15.9, +10.5]-0.7 points
Jev 1.13.0 (TypeSafe AI)73.0% (81/111)75.2% (82/109)-2.3 points [-13.8, +9.3]-0.7 points
system-one-open (Gemma 4 E2B LoRA on an L4)48.6% (54/111)49.5% (54/109)-0.9 points [-14.1, +12.3]-0.7 points
Bespoke Nimble 9B (Bespoke Labs)62.2% (69/111)68.8% (75/109)-6.6 points [-19.2, +5.9]-0.7 points
openjev-sglang (Qwen3.6-35B-A3B on SGLang)73.0% (81/111)69.7% (76/109)+3.2 points [-8.7, +15.2]-0.7 points
GPT-5.6 Luna (low reasoning effort)96.4% (107/111)92.7% (101/109)+3.7 points [-2.3, +9.7]-0.7 points
Gemini 3.1 Flash-Lite73.9% (82/111)76.1% (83/109)-2.3 points [-13.7, +9.2]-0.7 points
open-jev-deberta-v3-large (local CPU)37.8% (42/111)34.9% (38/109)+3.0 points [-9.7, +15.7]-0.7 points
DeepSeek V4.1 Flash (thinking default)96.4% (107/111)93.6% (102/109)+2.8 points [-2.9, +8.6]-0.7 points
Needle 3 (Cactus, 2-bit, local CPU) (partial)38.6% (17/44) (0/0) points [, ]-0.7 points
Qwen3.8 27B (Chutes TEE) (partial)92.2% (47/51) (0/0) points [, ]-0.7 points
Needle 3, options as tools (post-hoc adapter mode) (partial) (0/0) (0/0) points [, ]-0.7 points
open-alternative-jev (Qwen3.5-4B, IkerMoel)56.8% (63/111)56.9% (62/109)-0.1 points [-13.2, +13.0]-0.7 points
BAAI bge-reranker-v2-m337.8% (42/111)35.8% (39/109)+2.1 points [-10.7, +14.8]-0.7 points
Certo v1 (AltSlate Labs)33.3% (37/111)30.3% (33/109)+3.1 points [-9.2, +15.4]-0.7 points
classifier.dev (fast tier)70.3% (78/111)70.6% (77/109)-0.4 points [-12.4, +11.7]-0.7 points
decider-2b (Mapika)49.5% (55/111)45.0% (49/109)+4.6 points [-8.6, +17.8]-0.7 points
decider-35b-a3b (Mapika)66.7% (74/111)64.2% (70/109)+2.4 points [-10.1, +15.0]-0.7 points
decision-machine-1 (milliseconds.ai)48.6% (54/111)45.0% (49/109)+3.7 points [-9.5, +16.9]-0.7 points
djev (thinking)76.6% (85/111)78.9% (86/109)-2.3 points [-13.3, +8.7]-0.7 points
djev (Maisa, diffusion-gemma)67.6% (75/111)71.6% (78/109)-4.0 points [-16.1, +8.2]-0.7 points
GLiNER2 large (Fastino)36.9% (41/111)35.8% (39/109)+1.2 points [-11.6, +13.9]-0.7 points
GLiNER2.5 multi (Fastino, 287M)33.3% (37/111)42.2% (46/109)-8.9 points [-21.6, +3.9]-0.7 points
GLiNER2.5 small (Fastino, 74M)31.5% (35/111)34.9% (38/109)-3.3 points [-15.8, +9.1]-0.7 points
GLiNER2 (Fastino, gliner2.5-base)36.9% (41/111)35.8% (39/109)+1.2 points [-11.6, +13.9]-0.7 points
Alibaba GTE Reranker ModernBERT-base31.5% (35/111)35.8% (39/109)-4.2 points [-16.7, +8.2]-0.7 points
jeff (Logan Markewich, GLiFormer 400M)38.7% (43/111)36.7% (40/109)+2.0 points [-10.8, +14.8]-0.7 points
jev-local (Qwen3.5-9B)58.6% (65/111)59.6% (65/109)-1.1 points [-14.1, +11.9]-0.7 points
jqv (Qwen3-32B zero-shot)61.3% (68/111)67.9% (74/109)-6.6 points [-19.2, +6.0]-0.7 points
kev 0.5B29.7% (33/111)32.1% (35/109)-2.4 points [-14.6, +9.8]-0.7 points
kev 0.6B (research preview)43.2% (48/111)36.7% (40/109)+6.5 points [-6.4, +19.5]-0.7 points
kev 4B (research preview)36.9% (41/111)47.7% (52/109)-10.8 points [-23.8, +2.2]-0.7 points
kev 8B (research preview)45.0% (50/111)49.5% (54/109)-4.5 points [-17.7, +8.7]-0.7 points
Laya (Convai Innovations, ModernBERT-large 421M)35.1% (39/111)33.0% (36/109)+2.1 points [-10.4, +14.6]-0.7 points
LitJev (Qwen3.8-27B)72.1% (80/111)74.3% (81/109)-2.2 points [-13.9, +9.5]-0.7 points
Mixedbread mxbai-rerank-base-v236.0% (40/111)44.0% (48/109)-8.0 points [-20.9, +4.9]-0.7 points
Open-Jev 2B (Zefan Cai)41.4% (46/111)44.0% (48/109)-2.6 points [-15.7, +10.5]-0.7 points
Open-Jev 9B (Zefan Cai)59.5% (66/111)62.4% (68/109)-2.9 points [-15.8, +10.0]-0.7 points
OpenDecision (ModernBERT-large zero-shot)34.2% (38/111)32.1% (35/109)+2.1 points [-10.3, +14.6]-0.7 points
OpenJev (thinking, BF16)76.6% (85/111)79.8% (87/109)-3.2 points [-14.1, +7.7]-0.7 points
openJev Verdict 1.436.9% (41/111)38.5% (42/109)-1.6 points [-14.4, +11.2]-0.7 points
openJev Verdict (heman10x, ModernBERT-base 151M)37.8% (42/111)38.5% (42/109)-0.7 points [-13.5, +12.1]-0.7 points
Qwen3-Reranker-4B49.5% (55/111)50.5% (55/109)-0.9 points [-14.1, +12.3]-0.7 points
reflex-27b (Qwen3.8-27B)75.7% (84/111)76.1% (83/109)-0.5 points [-11.8, +10.8]-0.7 points
reflex 4B (kshetrajna12)60.4% (67/111)66.1% (72/109)-5.7 points [-18.4, +7.0]-0.7 points
SimpleJev Qwen3.6-35B-A3B65.8% (73/111)67.0% (73/109)-1.2 points [-13.7, +11.3]-0.7 points
SimpleJev Qwen3.8-27B73.9% (82/111)76.1% (83/109)-2.3 points [-13.7, +9.2]-0.7 points
smalljev semantic-v939.6% (44/111)36.7% (40/109)+2.9 points [-9.9, +15.8]-0.7 points
Winnow-12B Q873.0% (81/111)68.8% (75/109)+4.2 points [-7.8, +16.2]-0.7 points
ZeroEntropy zerank-251.4% (57/111)43.1% (47/109)+8.2 points [-4.9, +21.4]-0.7 points

Accuracy is correct / attempted; invalid responses count as incorrect. The interval is the unpooled two-sample normal 95% interval for a difference in proportions. Partial systems are shown but excluded from the field mean.

Public-split policy. Training on JevBench's public split is allowed and should be declared with each submission. Rankings continue to use all benchmark items. We report held-out results separately so that specialisation on public tasks is visible. Held-out means not publicly released, not guaranteed unseen: hosted systems receive these tasks during evaluation. We periodically issue fresh tasks to reduce the value of prior exposure.

Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole question — state, rubric and options — about 950 input tokens for Jev 1.13.0, so at its $0.042 per million input tokens 1,000 decisions cost $0.0399.

How costs are estimated

One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.

Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.

  • SemIf~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
  • Winnow-12B Q8~$0.037 est. per 1,000 decisions: OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • reflex 4B~$0.022 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • jqv~$0.056 est. per 1,000 decisions: OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • decider-35b-a3b~$0.067 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • open-alternative-jev~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
  • system-one-open~$0.015 est. per 1,000 decisions: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
  • OpenJev~$0.066 est. per 1,000 decisions: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
  • SimpleJev Qwen3.8-27B~$0.104 est. per 1,000 decisions: OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openjev-sglang~$0.131 est. per 1,000 decisions: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
  • reflex-27b~$0.181 est. per 1,000 decisions: OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • LitJev~$0.163 est. per 1,000 decisions: OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 0.6B~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • SimpleJev Qwen3.6-35B-A3B~$0.116 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • djev~$0.274 est. per 1,000 decisions: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures included
  • jev-local~$0.077 est. per 1,000 decisions: OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • decider-2b~$0.020 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Bespoke Nimble 9B~$0.166 est. per 1,000 decisions: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count))
  • OpenJev~$0.255 est. per 1,000 decisions: same hosted reference x 1778 billed input and 315 thought output tokens per decision
  • kev 4B~$0.019 est. per 1,000 decisions: DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 8B~$0.073 est. per 1,000 decisions: OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open-Jev 9B~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • system-one~$0.089 est. per 1,000 decisions: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
  • jeff~$0.0060 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Laya~$0.0029 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open-Jev 2B~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • OpenDecision~$0.0066 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openJev Verdict 1.4~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • openJev Verdict~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • kev 0.5B~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2 large~$0.0077 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • smalljev semantic-v9~$0.025 est. per 1,000 decisions: submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • open-jev-deberta-v3-large~$0.0073 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
  • GLiNER2.5 multi~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • GLiNER2.5 small~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • Certo v1~$0.0010 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • classifier.dev~$0.0033 est. per 1,000 decisions: ESTIMATE from the published paid plan (the free tier was used): classifier.dev Pro $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-19) = $0.0033 per 1,000 decisions at full use; one decision = one classification. Lower use costs more per decision: at a tenth of that allowance it is $0.033 per 1,000, and the free tier (20,000 fast classifications a day, which is what this run used) costs nothing.
  • Qwen3.8 27B~$2.669 est. per 1,000 decisions: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 416 input and 393 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision
  • Needle 3, options as tools~$0.014 est. per 1,000 decisions: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 383 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed. [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Needle 3~$0.024 est. per 1,000 decisions: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 383 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision
Reference prices by size class ($ per million input / output tokens)
  • dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33
  • dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55
  • dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15
  • encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005
  • generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201
  • moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3
  • moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9

Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).

Correction, v1.2.3 (20 September 2026): every price recomputed, each decision counted once

usd_per_1000_v11_tiers = 1000 x (mean input tokens per decision x $/M in + output tokens charged x $/M out) / 1e6, over all 314 v1.1 decisions (72 easy + 242 standard+judge), each decision counted exactly once and priced exactly once. A metered row uses the provider's own tariff and its own measured token counts, including the requests whose answer could not be parsed; an estimated row uses the reference tariff for its weights or size class and, when the run reports no usage, the input tokens of the gemini-3.1-flash-lite run on the same prompts over the same 314 decisions. usd_per_1000 = (v11 x 314 + hard x 220) / 534.

  • The v1.1 and v1.1.3 aggregations built their cost average from a row list that contained the 242-decision standard+judge run twice (once as the standard tier, once as the judge tier) and the 72 easy decisions once: 556 rows instead of 314. The standard and judge tiers were therefore over-weighted in the price, which made the affected rows look 1.5-3.3 % more expensive than they are.
  • Rows without their own token counts were priced at the input tokens of the gemini-3.1-flash-lite run measured on the 242 standard+judge decisions only (452 per decision) and that figure was applied to all 314 v1.1 decisions, which excludes the shorter easy tier. Over all 314 decisions the same run averages 383.41 input tokens, which is the figure used from v1.2.3 on. This made the affected rows look 4-11 % more expensive.
  • A metered row's price left out the requests whose answer came back unparseable. Those requests returned HTTP 200 with generated tokens and were billed, and JevBench already counts them as wrong answers, so from v1.2.3 they are priced too. Only DeepSeek V4.1 Flash had any (9 of its 314 v1.1 decisions); its price rises by 2.6 %.
  • No tariff was wrong. The hard-tier costs, and classifier.dev's flat plan price, were already correct.

No tariff, measurement, item, answer or rank changed. The prices before and after:

  • Jev 1.13.0 — $0.0406 → $0.0399 (-1.72 %)
  • SemIf — $0.0230 → $0.0224 (-2.31 %)
  • system-one-open — $0.0157 → $0.0149 (-5.13 %)
  • OpenJev — $0.0672 → $0.0656 (-2.36 %)
  • GPT-5.6 Luna — $0.2473 → $0.2419 (-2.17 %)
  • openjev-sglang — $0.1346 → $0.1313 (-2.48 %)
  • Bespoke Nimble 9B — $0.1085 → $0.1049 (-3.28 %)
  • Gemini 3.1 Flash-Lite — $0.2682 → $0.2638 (-1.65 %)
  • DeepSeek V4.1 Flash — $0.5788 → $0.5937 (+2.57 %)
  • system-one — $0.0915 → $0.0894 (-2.29 %)
  • openJev Verdict — $0.0039 → $0.0037 (-5.21 %)
  • GLiNER2 — $0.0039 → $0.0037 (-5.21 %)
  • open-jev-deberta-v3-large — $0.0077 → $0.0073 (-5.20 %)
  • Qwen3.8 27B — $2.7110 → $2.6691 (-1.55 %)
  • Needle 3, options as tools — $0.0162 → $0.0144 (-11.39 %)
  • Needle 3 — $0.0249 → $0.0238 (-4.36 %)

Who could not be measured, and why

An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank; so are the honorable mentions, which are complete runs that simply are not ranked.

  • open-jev (Dasein Labs)MLX on Apple Silicon only. Its own README says Linux containers cannot reach the Apple GPU, so a RunPod NVIDIA GPU cannot run it.
  • open-jev (JoshuaSP)A DiffusionGemma 26B-A4B serving wrapper rather than new trained weights. It was demonstrated on an H100; no public endpoint exists and no suitable 80 GB RunPod host was available in this round.
  • mini-jev (Mikhail Rakutko (r-ms))Public weights exist and fit a normal GPU, but the implementation covers Choice/Noul and explicitly does not measure Score. A faithful full-suite adapter would require new interface work rather than a mechanical endpoint adapter.
  • system-one-gemma (Akash Kamat)The adapter is public, but its Gemma base is gated behind Google’s licence terms. We do not accept binding terms on Florian’s behalf.
  • jevlike (Vincent Wang-Maścianica)Only Doom and chess vision checkpoints are released; there is no general text-decision checkpoint for this suite.
  • AlexWortega/openjev (Alex Wortega)Its released NLI and task-specific heads do not define a distribution over an arbitrary supplied label set. Inventing that mapping would measure our assumption.
  • Needle 3 (Cactus Compute)Its native response is a chosen label plus one accept/refuse confidence, not a categorical distribution over the supplied labels. The options-as-tools adaptation remains published as a partial run.
  • Succinct Router 14M (Pedro Marques)A router over three fixed GPT settings, not a general typed-decision model.
  • jev-model-router, Director, Loki (various)Applications built on decision models, not decision models themselves.
  • ProgramAsWeights (ProgramAsWeights)The compiler still requires GitHub authentication and the available path would expose held-out rubrics to a third party. No public weights or anonymous endpoint are available.
  • EigenJev (EigenJev)The endpoint requires authentication and no public weights or runnable implementation are published.
  • NanoJev (NanoJev)Public weights exist, but the server exposes a different schema (including boolean rather than Noul) and lacks the full structured/null contract. It needs substantive compatibility work before a fair full-suite run.
  • Werr (pCwOrM)Its documented server imports a module (scratch.jevbench_eval.optimize_werr_jevbench) that is not in the public repository, so the submitted configuration cannot be started; its engine also sends telemetry about each request to an outside server by default.
  • DIY Jev (VakeDomen)The repository named in the request (github.com/VakeDomen/DIY-Jev) answers 404, so there is nothing to run.
  • SimpleJev RWKV variants (SimpleJev)The public demo exposes RWKV IDs, but it does not identify their exact checkpoints or licences. Without reproducible model provenance, we do not publish benchmark rows for them.
Method and tiers

Geometric mean of Intelligence, Calibration, Speed and Cost, 25 % each. If chance-corrected Intelligence is below 50, multiply by (Intelligence / 50)^2; at or above 50 there is no penalty.

Revision v1.3.0. v1.3.0 scoring-only release: Intelligence is chance-corrected per tier and scores below 50 receive the growing near-chance penalty. Calibration, Speed, Cost, ranking eligibility, tasks and measurements are unchanged.

  • easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
  • standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
  • judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
  • hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.

Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.

A service running another entrant's model is listed, but not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number. Which rows this applies to, and why: Honorable mentions — services built on another entrant's model.

v1.2 numbers are not comparable with v1.1 or v1.0 (different tiers and scoring). The v1.0 page keeps its own numbers, calibration plots and per-family tables.

Limits
  • 534 decisions is a pilot, not a census, and it is English-only.
  • The weights are a choice. The JevBench Score weights the four axes equally and multiplies rather than adds them; if a wrong decision costs you more than a slow or expensive one, pick “Emphasis on Accuracy” above — the table of views shows what other weightings would do.
  • The latency adjustment (×2, +0.15 s on our own servers) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. The official Jev API is presumably under high load, given the public interest. Serving under load trades per-user speed for throughput: in the NVIDIA chart shown by SemiAnalysis, moving to the throughput-maximising setting cuts per-user tokens per second by far more than 2×. That chart is a 1.8T mixture-of-experts model on GPU clusters, not a 4B model on one GPU, so it supports the direction and size of the effect, not our exact factor. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo, and a measurement under load is planned.
  • Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.
  • Latency is one origin at one time of day; hosted endpoints, public demos and a local CPU are different kinds of latency. Public demo endpoints are shared with everyone else using them.
  • Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.
Credit

Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.

Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.