Tiny League

4 min read Original article ↗

Benchmark scores for recent small open-weight AI models.

Generated 2026-09-14T11:44:02Z

Changelog
  • 14 Sept 2026 — Added IFM/K2-Horizon-3.7B. All K2 Horizon models are now on Artificial Analysis: added the AA Intelligence Index (v4.3) to K2-Horizon-3.7B (16), K2-Horizon-7B (21) and K2-Horizon-MoVA-36B-A4B (26).
  • 13 Sept 2026 — Data-pipeline bug fixes: duplicate benchmark entries per model (several reasoning levels under one name) no longer hide data — the highest value is now shown and the disagreement is logged under conflicts; the Artificial Analysis fetcher no longer risks grabbing another model's score or dropping unrelated gaps; new models get vision support detected from the real config.json and are rejected when parameter metadata is missing.
  • 12 Sept 2026 — Added Agnes-AI/Agnes-3.0-Flash (33.1B dense, vision). Added a License column to the Models table.
  • 10 Sept 2026 — Added inclusionAI/Ling-3.0-flash-VL (124.8B MoE, vision) and meta-models/Muse-Glimmer-30B (29.8B dense, vision).
  • 9 Sept 2026 — Added nex-agi/Nex-N2.5-mini (35.1B MoE, vision). Added a model filter. Fixed the mobile layout (tables now scroll horizontally, heatmap keeps the model column pinned).
  • 8 Sept 2026 — Removed the Fara family (and JetBrains/Mellum2-12B); benchmark columns now show specific versions, grouped by family.
  • 7 Sept 2026 — Added openbmb/MiniCPM5-2B (2.5B dense), IFM K2-Horizon-MoVA-36B-A4B (36B MoE, 4B active) and K2-Horizon-7B (9B dense).
  • 4 Sept 2026 — Initial release with 26 open-weight models.

Models

Listed by release date, newest first. Click any column header to sort (Params / Benchmarks / AA-Index sort numerically, Released by date; click again to flip direction; missing values sink to the bottom). A V badge marks models that support vision / multimodal (image) input.

Coverage heatmap

Cell shows the score for that model×benchmark; grey = absent / no numeric score. Cell colour is scaled per benchmark (light → deep blue = lower → higher); the gold cell is the best model on that benchmark. Columns are ordered by how many models report each benchmark. Click a benchmark column to rank models by it; click the Model column to sort by name.

lowerhigherbest on benchmarkVsupports vision / multimodalMoEMixture-of-Experts (MoE)

ModelTerminal-Bench
2.1
Terminal-Bench
2.0
Terminal-Bench
Hard
AA-Intelligence-…GPQA DiamondHLESWE-bench ProSWE-bench Verifi…IFBenchSciCodeSWE-bench Multil…AIME
2026
AIME
2025
AA-LCRHMMT
Feb 2026
HMMT
Feb 2025
LiveCodeBench
v6
LiveCodeBenchLiveCodeBench
2408-2505
τ³-benchBrowseCompBFCL
v4
BFCL
v3
MCP-AtlasMMLU-ProIMO AnswerBenchτ²-benchDeepSWE
1.1
DeepSWEMRCRMRCR
v2
MathVisionOmniDocBench 1.5CharXivGDPval-AA-V2IFEvalMulti-IFNL2Repo-BenchToolathlonVision2WebAA-Omniscience A…ClawEvalClawEval-MMCodeforces ELOGDPvalGDPval-AAJobBenchMATH500MMMU ProOSWorld-VerifiedWideSearchAgents' Last Exa…AndroidWorldArena-Hard-V2ArtifactsBenchBigBench Extra H…BirdBenchBrowseComp-zhClaw-GymCoWorkBenchERQALIFEbenchLongBench-v2MMLU-ProX liteMMMLUMedXPertQA MMOSWorldPinchBenchProfBenchRULERRealWorldQARecreationBenchSWE AtlasSWE-MMSkillsBenchToolathlon-Verif…WebArena-Verifie…
Qwen3.8-27BV73··3489.230.861.7·79.5·······90.3··········42.2···94.691.190.2···42.3·62.9··57.4···33.4··84.3·42.981.9······70.765.5·········85.947.1·38.6··64.8
Qwen3.8-Flash-NextVMoE···4291.735.962.5·81.3·81·····91.9··········58.7···95.7·90.6···48.173.564··64.4···55.7····51.284.5······73.972.3·····52.3···88.549.9·····
Qwen-AgentWorld-35B-A3BVMoE·············································································
gemma-4-12B-itV···1478.85.2·····77.5····72·······77.2·69···43.479.70.2··········1659····69.1······53········83.448.7···········
diffusiongemma-26B-A4B-itVMoE···1073.211.9·····69.1····69.1·······77.6·56.2···3270.50.3··········1429····54.3······47.6········81.549···········
DeepSeek-V4-Flash-DSparkMoE·56.9·3588.145.152.679··73.3···94.8··91.6··73.2··6986.288.4···78.7········47.8····3052·1395·······························
Laguna-S-2.1MoE70.2·····59.4···78.5·················40.4·········49.7·································46.2····
Laguna-XS-2.1MoE·37.5····47.670.9··63.1··································································
LFM2.5-8B-A1BMoE···7····56.5··5042.5········49.764.8···88.1········91.879.9···8.7······88.8·····························
Nanbeige4.2-3B·44.1··87.417.846.963.654.635.6···58.782.8·72.5······57.8·67.3···············52.2··74.3·············65··················
KAT-Coder-V2.5-DevV41.0·····46.069.4·44.263························································93.4·········
North-Mini-Code-1.0MoE·3631.11375.711.140.267.657.538.2······70.3····························································
granite-4.2-30b29.2··1566.4·33.35777.238.841.9·89.2··89.275.8··62·61.4··77.6···················1225········67.9··41.9······66.6····42.990.0·······
granite-4.2-8b20.6··1264.1·19.147.779.336.130.8·86.7··78.373.2··58.1·52.4··74.0···················1189········65.2··41.1······61.1····41.281.0·······
Hy-MT2-30B-A3BMoE········50.7··························89.8·········································
LongCat-Flash-Lite-Sparse·33.7·1169.5·40.668.2··59.365.7·4841.5·····48.6··45.679.249.496.0··44.7·················96.8·········61.9····53.6··············
Nemotron-3-Nano-Omni-30B-A3B-ReasoningVMoE···10······························································47.4··········
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16MoE24.6··1475.411.7·51.671.932.639.3··52·····9.337.0···81.9·········832································85.4·········
Ornith-1.5-9BV47···86.430.547.570.6··54.4·········56.4··54.2·············32.4···66.5········59.5························41.2·
Ornith-1.5-35B-A3BVMoE68.5···89.233.459.679··71.4·········67.6··70.2····22········46.248.7··72.5········67.8·····················39.8····
Ling-3.0-tinyMoE27.7··1273.49.3··63.624.2···58.770.3····20.8·62.7···71.0··········83.2········772········47.9······62.3···············
Ling-3.0-flashMoE57··25·22.756.6·74.5·72.493.2·65.187···82.8288273·65.5·83.7···90.8······87.7········1107····73.6···77······77.3············44.8··
Spark-X2.5-4B····67.412.344.441.67534.753.390.7·56.381.2····30.440.965.1·54.6·74.275.1········93·········································
K2-Horizon-MoVA-36B-A4BMoE58.6··2680.825.2···38.9···66.3·····26.8····················18.8····································
K2-Horizon-7B39.1··21·18.6·70.6·31.6···6873.3····25.859························································
MiniCPM5-2B8.6··1470.28.914.446.466.326.3·86.586.55963.8·69.1··20.839.766.6··70.8·97.1·······19.686.771.8··········94.6·········43.559.2···43.7··············
Nex-N2.5-miniVMoE73.4·····43.8·············83.4······36.1······1446····52.9······28.5··71.2·······················25.5·54.663.4
Ling-3.0-flash-VLVMoE···25···························84.991.381.3·····57.7··59.9··································
Muse-Glimmer-30BV51.7··1883.52251.2767743.6·94.7·80·····23.5···75.5········75.878.8953·············7465.9························44.3··
Agnes-3.0-FlashV···3685.0···74.238.1···68.3··························23····································
K2-Horizon-3.7B25.1··1665.412.9·68.6·25.9····70.5····17.7·50.9·······················································

AA-Intelligence-Index-v4.3 vs parameters

Comparisons

Bar charts compare the top 10 models (by score) that report the same benchmark (≥2 models within the parameter range above). Hover a bar for the exact value.

Per-benchmark detail

Terminal-Bench 2.1 (16 models)

Terminal-Bench 2.0 (5 models)

Terminal-Bench Hard (1 models)

AA-Intelligence-Index-v4.3 (21 models)

GPQA Diamond (20 models)

HLE (18 models)

SWE-bench Pro (18 models)

SWE-bench Verified (16 models)

IFBench (15 models)

SciCode (14 models)

SWE-bench Multilingual (13 models)

AIME 2026 (8 models)

AIME 2025 (4 models)

AA-LCR (11 models)

HMMT Feb 2026 (9 models)

HMMT Feb 2025 (2 models)

LiveCodeBench v6 (9 models)

LiveCodeBench (1 models)

LiveCodeBench 2408-2505 (1 models)

τ³-bench (11 models)

BrowseComp (10 models)

BFCL v4 (8 models)

BFCL v3 (1 models)

MCP-Atlas (8 models)

MMLU-Pro (8 models)

IMO AnswerBench (6 models)

τ²-bench (6 models)

DeepSWE 1.1 (3 models)

DeepSWE (2 models)

MRCR (3 models)

MRCR v2 (2 models)

MathVision (5 models)

OmniDocBench 1.5 (5 models)

CharXiv (4 models)

GDPval-AA-V2 (4 models)

IFEval (4 models)

Multi-IF (4 models)

NL2Repo-Bench (4 models)

Toolathlon (4 models)

Vision2Web (4 models)

AA-Omniscience Accuracy (3 models)

ClawEval (3 models)

ClawEval-MM (3 models)

Codeforces ELO (3 models)

GDPval (3 models)

GDPval-AA (3 models)

JobBench (3 models)

MATH500 (3 models)

MMMU Pro (3 models)

OSWorld-Verified (3 models)

WideSearch (3 models)

Agents' Last Exam (2 models)

AndroidWorld (2 models)

Arena-Hard-V2 (2 models)

ArtifactsBench (2 models)

BigBench Extra Hard (2 models)

BirdBench (2 models)

BrowseComp-zh (2 models)

Claw-Gym (2 models)

CoWorkBench (2 models)

ERQA (2 models)

LIFEbench (2 models)

LongBench-v2 (2 models)

MMLU-ProX lite (2 models)

MMMLU (2 models)

MedXPertQA MM (2 models)

OSWorld (2 models)

PinchBench (2 models)

ProfBench (2 models)

RULER (2 models)

RealWorldQA (2 models)

RecreationBench (2 models)

SWE Atlas (2 models)

SWE-MM (2 models)

SkillsBench (2 models)

Toolathlon-Verified (2 models)

WebArena-Verified (2 models)