AA-Briefcase: Agentic Knowledge Work Benchmark | Artificial Analysis

Artificial Analysis

6 min read Original article ↗

Results

AA-Briefcase Elo

AA-Briefcase is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better

Reasoning models are indicated by a lightbulb icon

AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.

Cost

AA-Briefcase Cost per Task

Mean cost (USD) per task to run AA-Briefcase, calculated from token usage and model pricing including representative cache hit rates

Reasoning models are indicated by a lightbulb icon

The total cost to run AA-Briefcase divided by the number of tasks (91 for full submission of tasks). Cost is calculated from token usage and model pricing, split across input, cache hit, cache write, reasoning, and answer token prices, including representative cache hit rates.

Example Task, Submissions, and Grading

Explore a representative AA-Briefcase week from the public Due Diligence scenario available via

Hugging Face.

The outputs and grading shown here illustrate what AA-Briefcase evaluates. Scores are shown for a

representative model set. Submissions and verdicts in this representative scenario do not contribute to a model's AA-Briefcase Elo or other benchmark scores.

Score Comparisons

AA-Briefcase Elo vs. Artificial Analysis Intelligence Index

AA-Briefcase Elo · Artificial Analysis Intelligence Index

AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.

Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

File Type Results

AA-Briefcase performance broken out by the file type of the deliverable (Excel, PowerPoint, PDF, Word, Other).

AA-Briefcase Rubric Pass Rate by File Type (Normalized)

Rubric pass rate by deliverable file type · Scores are normalized per file type across all models tested, where green represents the highest score for that file type and red represents the lowest score for that file type

Reasoning models are indicated by a lightbulb icon

File types are categorized by the required submission format, with “Other” covering formats such as HTML and LaTeX.

The share of binary rubric checks the submission passed (passed checks divided by total checks), aggregated across all AA-Briefcase tasks. Rubric checks are pass/fail criteria covering whether the deliverable includes required content and cites sources correctly, and whether it resolves planted cross-source conflicts.

Token Usage

AA-Briefcase Output Tokens per Task

Mean reasoning and answer tokens consumed per AA-Briefcase task

Reasoning models are indicated by a lightbulb icon

The number of output tokens used to run the evaluation, including visible answer tokens and reasoning tokens where reported by reasoning models.

Speed

Time per Task

Wall-clock time (minutes) per task: answer and reasoning generation plus tool execution time · Lower is better

Reasoning models are indicated by a lightbulb icon

Estimated wall-clock time per task: the sum of answer and reasoning tokens per task divided by the model’s canonical answer output speed, plus mean tool execution time per task. Lower is better.

Turns

Mean Turns per Task

Average number of model turns per AA-Briefcase task · Lower is better

Reasoning models are indicated by a lightbulb icon

This chart shows the average number of turns the agent takes per task. It is a rough proxy for how many actions, tool calls, and iteration cycles an agent is using to complete benchmark tasks.

Tool Usage

Tool invocations issued by each agent during AA-Briefcase: counts by tool category, mean tool calls per turn, and source-pool exploration coverage.

AA-Briefcase Tool Calls Breakdown, Avg per Task

Average tool invocations per AA-Briefcase task, bucketed by intent

Reasoning models are indicated by a lightbulb icon

Agent tool calls are grouped into six categories: explore (navigating and searching the workspace), read (reading file contents), write (creating or editing files), compute (running code or calculations), view image (visual inspection of files), and other (anything else).

Model Size (Open Weights Models Only)

AA-Briefcase Elo vs. Total Parameters

AA-Briefcase Elo · Size in parameters (billions) · Open weights models only

AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.

The total number of trainable weights and biases in the model, expressed in billions. These parameters are learned during training and determine the model's ability to process and generate responses.

Score vs. Release Date

AA-Briefcase Elo vs. Release Date

AA-Briefcase Elo · Model release date

AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.

Leaderboard

Creator

Name

Elo

CI

Release Date

1

Anthropic logoAnthropic

Claude Opus 5 (Adaptive Reasoning, Max Effort)1715-11 / +13Jul 2026
2

Anthropic logoAnthropic

Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)1690-11 / +12Jul 2026
3

Anthropic logoAnthropic

Claude Opus 5 (Adaptive Reasoning, High Effort)1606-11 / +12Jul 2026
4

SpaceXAI logoSpaceXAI

Grok 4.6 (high)1577-11 / +11Aug 2026
5

Anthropic logoAnthropic

Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)1574-10 / +10Jun 2026
6

Kimi logoKimi

Kimi K3 (max)1541-10 / +11Jul 2026
7

OpenAI logoOpenAI

GPT-5.6 Sol (max)1504-9 / +10Jul 2026
8

Anthropic logoAnthropic

Claude Opus 5 (Adaptive Reasoning, Medium Effort)1469-10 / +11Jul 2026
9

Alibaba logoAlibaba

Qwen3.8 Max1420-11 / +12Aug 2026
10

Anthropic logoAnthropic

Claude Sonnet 5 (Adaptive Reasoning, Max Effort)1384-9 / +10Jun 2026
11

Meta logoMeta

Muse Spark 1.2 (xhigh)1358-11 / +12Aug 2026
12

Anthropic logoAnthropic

Claude Opus 4.8 (Adaptive Reasoning, Max Effort)1340-9 / +8May 2026
13

SpaceXAI logoSpaceXAI

Grok 4.5 (high)1313-10 / +10Jul 2026
14

Anthropic logoAnthropic

Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)1292-10 / +9Jun 2026
15

DeepSeek logoDeepSeek

DeepSeek V4 Flash 0731 (Reasoning, Max Effort)1286-9 / +9Jul 2026
16

Anthropic logoAnthropic

Claude Opus 4.7 (Adaptive Reasoning, Max Effort)1276-8 / +9Apr 2026
17

Z AI logoZ AI

GLM-5.2 (max)1252-9 / +8Jun 2026
18

Anthropic logoAnthropic

Claude Opus 5 (Adaptive Reasoning, Low Effort)1225-10 / +9Jul 2026
19

Anthropic logoAnthropic

Claude Sonnet 5 (Adaptive Reasoning, High Effort)1193-9 / +9Jun 2026
20

OpenAI logoOpenAI

GPT-5.5 (xhigh)1150-8 / +8Apr 2026
21

Google logoGoogle

Gemini 3.7 Flash (high)1132-10 / +11Aug 2026
22

MiniMax logoMiniMax

MiniMax-M31107-7 / +9Jun 2026
23

OpenAI logoOpenAI

GPT-5.5 (high)1099-8 / +8Apr 2026
24

Anthropic logoAnthropic

Claude Opus 4.7 (Non-reasoning, High Effort)1083-8 / +8Apr 2026
25

Anthropic logoAnthropic

Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)1075-8 / +8Feb 2026
26

Anthropic logoAnthropic

Claude Sonnet 5 (Adaptive Reasoning, Medium Effort)1057-9 / +8Jun 2026
27

OpenAI logoOpenAI

GPT-5.5 (medium)1000-0 / +0Apr 2026
28

Z AI logoZ AI

GLM-5.1 (Reasoning)971-8 / +8Apr 2026
29

Google logoGoogle

Gemini 3.6 Flash (high)963-9 / +10Jul 2026
30

Anthropic logoAnthropic

Claude Sonnet 5 (Adaptive Reasoning, Low Effort)931-9 / +8Jun 2026
31

DeepSeek logoDeepSeek

DeepSeek V4 Pro (Reasoning, Max Effort)930-8 / +8Apr 2026
32

Thinking Machines logoThinking Machines

Inkling Small917-11 / +10Jul 2026
33

Alibaba logoAlibaba

Qwen3.7 Max914-8 / +8May 2026
34

Xiaomi logoXiaomi

MiMo-V2.5-Pro880-8 / +8Apr 2026
35

NVIDIA logoNVIDIA

Nemotron 3 Ultra 550B A55B (Reasoning)873-8 / +9Jun 2026
36

Google logoGoogle

Gemini 3.5 Flash (high)872-8 / +8May 2026
37

Google logoGoogle

Gemini 3.5 Flash (medium)871-9 / +8May 2026
38

OpenAI logoOpenAI

GPT-5.3 Codex (xhigh)870-8 / +8Feb 2026
39

Meta logoMeta

Muse Spark 1.1 (xhigh)869-10 / +11Jul 2026
40

Thinking Machines logoThinking Machines

Inkling (xhigh)842-10 / +10Jul 2026
41

DeepSeek logoDeepSeek

DeepSeek V4 Flash (Reasoning, Max Effort)833-9 / +8Apr 2026
42

Kimi logoKimi

Kimi K2.6819-9 / +8Apr 2026
43

Alibaba logoAlibaba

Qwen3.6 27B (Reasoning)810-10 / +9Apr 2026
44

SpaceXAI logoSpaceXAI

Grok 4.3 (high)760-9 / +8Apr 2026
45

OpenAI logoOpenAI

GPT-5.4 mini (xhigh)717-9 / +9Mar 2026
46

Meta logoMeta

Muse Spark643-10 / +10Apr 2026
47

Google logoGoogle

Gemini 3.5 Flash-Lite635-12 / +11Jul 2026
48

Anthropic logoAnthropic

Claude 4.5 Haiku (Reasoning)612-10 / +9Oct 2025
49

KwaiKAT logoKwaiKAT

KAT-Coder-Pro V1599-11 / +11Nov 2025
50

Alibaba logoAlibaba

Qwen3.5 397B A17B (Reasoning)554-11 / +10Feb 2026
51

Mistral logoMistral

Mistral Medium 3.5517-10 / +9Apr 2026
52

Google logoGoogle

Gemini 3.1 Pro Preview458-12 / +10Feb 2026
53

Google logoGoogle

Gemma 4 31B (Reasoning)374-12 / +11Apr 2026
54

Cohere logoCohere

Command A+369-15 / +13May 2026
55

Cohere logoCohere

North Mini Code239-15 / +14Jun 2026
56

Google logoGoogle

Gemini 3.1 Flash-Lite231-14 / +12Mar 2026
57

Upstage logoUpstage

Solar Pro 3138-14 / +12Apr 2026
58

MBZUAI Institute of Foundation Models logoMBZUAI Institute of Foundation Models

K2 Think V260-15 / +13Dec 2025
59

OpenAI logoOpenAI

gpt-oss-120b (high)8-8 / +13Aug 2025
60

OpenAI logoOpenAI

gpt-oss-20b (high)0-0 / +0Aug 2025
61

Meta logoMeta

Llama 4 Maverick0-0 / +0Apr 2025
62

NVIDIA logoNVIDIA

Nemotron 3 Super 120B A12B (Reasoning)0-0 / +0Mar 2026