Terminal-Bench 2.1
One harness to unlock the potential of all models
Compare Ante runs across models on the same Terminal-Bench 2.1 task set, using consistent parameters and verified benchmark results.
Leading Model
DeepSeek V4 Flash 0731max
Task Set89 tasks
Trials368 passed / 445 trials
We benchmark what we ship.Every eval uses a pinned public . No eval-only branches or benchmark-specific prompts.
The runs are auditable.Every result links its raw Harbor run, so anyone can inspect the trials behind the number.
The constraints are official.All trials follow the official Terminal-Bench parameters: 89 tasks, 5 trials per task, strict timeouts, and hardware limits.
Model org
All model orgsTB 2.1 · Same parameters: 89 tasks · 5 trials/task · Updated Aug 9, 2026