Optima | Artificial Analysis

3 min read Original article ↗

Optima

Build your own custom benchmark

Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom benchmarks around your own tasks, so you can compare models on performance, cost, and time efficiency

Why build your benchmark with Optima

Results specific to your use case

Build a benchmark from your own data and tasks, and get results on the latest models as soon as they're released.

Cut cost and time by over 10x

Compare on more than performance, with every result showing cost and time per task efficiency comparisons across models

Built on Artificial Analysis grading expertise

Grade against custom rubric criteria, or use our panel of judges from major evaluations to rank models head-to-head

How it works

1

Give Optima your context

Describe the work, attach a few examples, and the build agent drafts the tasks and rubrics with you.

Describe your use case|Attach examplespptxAA-Merch_Sales_Pitch_DeckpptxAA-Merch_Startup_Swag_Pitchtwo decks we were happy with — this is what good looks likeBuild agentdrafts tasks + rubricsYour drafted benchmarkSales Deck Creation5 tasks · 10 criteria · head-to-headNorthstar Hotels welcome-kit pilotCrescent Arts Museum collection pitchApex Trails 12-month merch programmeTidepool summer pre-season launchCity Sound Festival activation

2

Choose your evaluation type

Q&AObjective grading

Example task prompt

Which HS tariff code applies to lithium-ion e-bike batteries?

How it's graded

The rubric carries the expected answer. A judge model checks each response against it, so the score is deterministic — right or wrong, criterion by criterion.

3

Run across models

Your benchmark24 tasksFlag the risky clausesExtract the payment termsDraft the counterparty summarysame tasks, same conditions, every modelRunsandbox + toolsModels you pickedClaude Fable 5GPT-5.6 SolKimi K3Gemini 3.6 Flashtokens, cost and time recorded per task24/24

Or bring your own agent

Your own agent can compete in the same run, over HTTP.

Claude Opus 4.6GPT-5.2Gemini 3 Proyour-agentPOST /runs/9f2c/artifactsBenchmarkrun24 tasks · same judgesLeaderboard1Claude Opus 4.63GPT-5.24Gemini 3 Pro2your-agent

4

Grade and decide

strong and cheapScoreCost per task$0.00$0.05$0.10$0.15$0.20Claude Fable 560 · 74s per taskGPT-5.6 Sol59 · 52s per taskKimi K357 · 61s per taskGemini 3.6 Flash50 · 29s per task

Pricing

Building and running a benchmark is priced on token usage. Grading is priced per unit judged.

  • Creating and running a benchmark is charged at the raw token cost of the models used, with nothing added on top.
  • Rubric grading is $0.002 per criterion, per model with standard judges, or $0.040 with premium judges.
  • Pairwise grading is $0.006 per match with standard judges, or $0.150 with premium judges.
  • Rubric grading uses one judge from the selected tier; pairwise grading uses the full judge panel. Standard uses high-capability models, while premium uses frontier models that cost several times more to run.

At the start of benchmark creation and each benchmark run, we hold an amount of your credit based on our cost estimate for that stage. At the end of the stage you are only ever charged for the tokens actually used, regardless of the estimate.

Grading works the other way round: the rates above are a price, not an estimate. We hold the quoted total for the whole pass, then charge the rate for each criterion or match actually judged — so a grading pass never costs more than its quote, and costs less when it judges fewer units than planned.

Find the best model for your work

Bring your own tasks or describe your use case, and get graded results with the costs attached.

Try Optima