Results
AA-Briefcase Elo
AA-Briefcase is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better
Reasoning models are indicated by a lightbulb icon
AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.
Cost
AA-Briefcase Cost per Task
Mean cost (USD) per task to run AA-Briefcase, calculated from token usage and model pricing including representative cache hit rates
Reasoning models are indicated by a lightbulb icon
The total cost to run AA-Briefcase divided by the number of tasks (91 for full submission of tasks). Cost is calculated from token usage and model pricing, split across input, cache hit, cache write, reasoning, and answer token prices, including representative cache hit rates.
Example Task, Submissions, and Grading
Explore a representative AA-Briefcase week from the public Due Diligence scenario available via
The outputs and grading shown here illustrate what AA-Briefcase evaluates. Scores are shown for a
representative model set. Submissions and verdicts in this representative scenario do not contribute to a model's AA-Briefcase Elo or other benchmark scores.
Score Comparisons
AA-Briefcase Elo vs. Artificial Analysis Intelligence Index
AA-Briefcase Elo · Artificial Analysis Intelligence Index
AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.
Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
File Type Results
AA-Briefcase performance broken out by the file type of the deliverable (Excel, PowerPoint, PDF, Word, Other).
AA-Briefcase Rubric Pass Rate by File Type (Normalized)
Rubric pass rate by deliverable file type · Scores are normalized per file type across all models tested, where green represents the highest score for that file type and red represents the lowest score for that file type
Reasoning models are indicated by a lightbulb icon
File types are categorized by the required submission format, with “Other” covering formats such as HTML and LaTeX.
The share of binary rubric checks the submission passed (passed checks divided by total checks), aggregated across all AA-Briefcase tasks. Rubric checks are pass/fail criteria covering whether the deliverable includes required content and cites sources correctly, and whether it resolves planted cross-source conflicts.
Token Usage
AA-Briefcase Output Tokens per Task
Mean reasoning and answer tokens consumed per AA-Briefcase task
Reasoning models are indicated by a lightbulb icon
The number of output tokens used to run the evaluation, including visible answer tokens and reasoning tokens where reported by reasoning models.
Speed
Time per Task
Wall-clock time (minutes) per task: answer and reasoning generation plus tool execution time · Lower is better
Reasoning models are indicated by a lightbulb icon
Estimated wall-clock time per task: the sum of answer and reasoning tokens per task divided by the model’s canonical answer output speed, plus mean tool execution time per task. Lower is better.
Turns
Mean Turns per Task
Average number of model turns per AA-Briefcase task · Lower is better
Reasoning models are indicated by a lightbulb icon
This chart shows the average number of turns the agent takes per task. It is a rough proxy for how many actions, tool calls, and iteration cycles an agent is using to complete benchmark tasks.
Tool Usage
Tool invocations issued by each agent during AA-Briefcase: counts by tool category, mean tool calls per turn, and source-pool exploration coverage.
AA-Briefcase Tool Calls Breakdown, Avg per Task
Average tool invocations per AA-Briefcase task, bucketed by intent
Reasoning models are indicated by a lightbulb icon
Agent tool calls are grouped into six categories: explore (navigating and searching the workspace), read (reading file contents), write (creating or editing files), compute (running code or calculations), view image (visual inspection of files), and other (anything else).
Model Size (Open Weights Models Only)
AA-Briefcase Elo vs. Total Parameters
AA-Briefcase Elo · Size in parameters (billions) · Open weights models only
AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.
The total number of trainable weights and biases in the model, expressed in billions. These parameters are learned during training and determine the model's ability to process and generate responses.
Score vs. Release Date
AA-Briefcase Elo vs. Release Date
AA-Briefcase Elo · Model release date
AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.
Leaderboard
Creator | Name | Elo | CI | Release Date | |
|---|---|---|---|---|---|
| 1 |
| Claude Opus 5 (Adaptive Reasoning, Max Effort) | 1715 | -11 / +13 | Jul 2026 |
| 2 |
| Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | 1690 | -11 / +12 | Jul 2026 |
| 3 |
| Claude Opus 5 (Adaptive Reasoning, High Effort) | 1606 | -11 / +12 | Jul 2026 |
| 4 |
| Grok 4.6 (high) | 1577 | -11 / +11 | Aug 2026 |
| 5 |
| Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | 1574 | -10 / +10 | Jun 2026 |
| 6 |
| Kimi K3 (max) | 1541 | -10 / +11 | Jul 2026 |
| 7 |
| GPT-5.6 Sol (max) | 1504 | -9 / +10 | Jul 2026 |
| 8 |
| Claude Opus 5 (Adaptive Reasoning, Medium Effort) | 1469 | -10 / +11 | Jul 2026 |
| 9 |
| Qwen3.8 Max | 1420 | -11 / +12 | Aug 2026 |
| 10 |
| Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | 1384 | -9 / +10 | Jun 2026 |
| 11 |
| Muse Spark 1.2 (xhigh) | 1358 | -11 / +12 | Aug 2026 |
| 12 |
| Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | 1340 | -9 / +8 | May 2026 |
| 13 |
| Grok 4.5 (high) | 1313 | -10 / +10 | Jul 2026 |
| 14 |
| Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | 1292 | -10 / +9 | Jun 2026 |
| 15 |
| DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | 1286 | -9 / +9 | Jul 2026 |
| 16 |
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | 1276 | -8 / +9 | Apr 2026 |
| 17 |
| GLM-5.2 (max) | 1252 | -9 / +8 | Jun 2026 |
| 18 |
| Claude Opus 5 (Adaptive Reasoning, Low Effort) | 1225 | -10 / +9 | Jul 2026 |
| 19 |
| Claude Sonnet 5 (Adaptive Reasoning, High Effort) | 1193 | -9 / +9 | Jun 2026 |
| 20 |
| GPT-5.5 (xhigh) | 1150 | -8 / +8 | Apr 2026 |
| 21 |
| Gemini 3.7 Flash (high) | 1132 | -10 / +11 | Aug 2026 |
| 22 |
| MiniMax-M3 | 1107 | -7 / +9 | Jun 2026 |
| 23 |
| GPT-5.5 (high) | 1099 | -8 / +8 | Apr 2026 |
| 24 |
| Claude Opus 4.7 (Non-reasoning, High Effort) | 1083 | -8 / +8 | Apr 2026 |
| 25 |
| Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | 1075 | -8 / +8 | Feb 2026 |
| 26 |
| Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) | 1057 | -9 / +8 | Jun 2026 |
| 27 |
| GPT-5.5 (medium) | 1000 | -0 / +0 | Apr 2026 |
| 28 |
| GLM-5.1 (Reasoning) | 971 | -8 / +8 | Apr 2026 |
| 29 |
| Gemini 3.6 Flash (high) | 963 | -9 / +10 | Jul 2026 |
| 30 |
| Claude Sonnet 5 (Adaptive Reasoning, Low Effort) | 931 | -9 / +8 | Jun 2026 |
| 31 |
| DeepSeek V4 Pro (Reasoning, Max Effort) | 930 | -8 / +8 | Apr 2026 |
| 32 |
| Inkling Small | 917 | -11 / +10 | Jul 2026 |
| 33 |
| Qwen3.7 Max | 914 | -8 / +8 | May 2026 |
| 34 |
| MiMo-V2.5-Pro | 880 | -8 / +8 | Apr 2026 |
| 35 |
| Nemotron 3 Ultra 550B A55B (Reasoning) | 873 | -8 / +9 | Jun 2026 |
| 36 |
| Gemini 3.5 Flash (high) | 872 | -8 / +8 | May 2026 |
| 37 |
| Gemini 3.5 Flash (medium) | 871 | -9 / +8 | May 2026 |
| 38 |
| GPT-5.3 Codex (xhigh) | 870 | -8 / +8 | Feb 2026 |
| 39 |
| Muse Spark 1.1 (xhigh) | 869 | -10 / +11 | Jul 2026 |
| 40 |
| Inkling (xhigh) | 842 | -10 / +10 | Jul 2026 |
| 41 |
| DeepSeek V4 Flash (Reasoning, Max Effort) | 833 | -9 / +8 | Apr 2026 |
| 42 |
| Kimi K2.6 | 819 | -9 / +8 | Apr 2026 |
| 43 |
| Qwen3.6 27B (Reasoning) | 810 | -10 / +9 | Apr 2026 |
| 44 |
| Grok 4.3 (high) | 760 | -9 / +8 | Apr 2026 |
| 45 |
| GPT-5.4 mini (xhigh) | 717 | -9 / +9 | Mar 2026 |
| 46 |
| Muse Spark | 643 | -10 / +10 | Apr 2026 |
| 47 |
| Gemini 3.5 Flash-Lite | 635 | -12 / +11 | Jul 2026 |
| 48 |
| Claude 4.5 Haiku (Reasoning) | 612 | -10 / +9 | Oct 2025 |
| 49 |
| KAT-Coder-Pro V1 | 599 | -11 / +11 | Nov 2025 |
| 50 |
| Qwen3.5 397B A17B (Reasoning) | 554 | -11 / +10 | Feb 2026 |
| 51 |
| Mistral Medium 3.5 | 517 | -10 / +9 | Apr 2026 |
| 52 |
| Gemini 3.1 Pro Preview | 458 | -12 / +10 | Feb 2026 |
| 53 |
| Gemma 4 31B (Reasoning) | 374 | -12 / +11 | Apr 2026 |
| 54 |
| Command A+ | 369 | -15 / +13 | May 2026 |
| 55 |
| North Mini Code | 239 | -15 / +14 | Jun 2026 |
| 56 |
| Gemini 3.1 Flash-Lite | 231 | -14 / +12 | Mar 2026 |
| 57 |
| Solar Pro 3 | 138 | -14 / +12 | Apr 2026 |
| 58 |
| K2 Think V2 | 60 | -15 / +13 | Dec 2025 |
| 59 |
| gpt-oss-120b (high) | 8 | -8 / +13 | Aug 2025 |
| 60 |
| gpt-oss-20b (high) | 0 | -0 / +0 | Aug 2025 |
| 61 |
| Llama 4 Maverick | 0 | -0 / +0 | Apr 2025 |
| 62 |
| Nemotron 3 Super 120B A12B (Reasoning) | 0 | -0 / +0 | Mar 2026 |