Optima
Build your own custom benchmark
Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom benchmarks around your own tasks, so you can compare models on performance, cost, and time efficiency
Why build your benchmark with Optima
Results specific to your use case
Build a benchmark from your own data and tasks, and get results on the latest models as soon as they're released.
Cut cost and time by over 10x
Compare on more than performance, with every result showing cost and time per task efficiency comparisons across models
Built on Artificial Analysis grading expertise
Grade against custom rubric criteria, or use our panel of judges from major evaluations to rank models head-to-head
How it works
1
Give Optima your context
Describe the work, attach a few examples, and the build agent drafts the tasks and rubrics with you.
2
Choose your evaluation type
Q&AObjective grading
Example task prompt
Which HS tariff code applies to lithium-ion e-bike batteries?
How it's graded
The rubric carries the expected answer. A judge model checks each response against it, so the score is deterministic — right or wrong, criterion by criterion.
3
Run across models
Or bring your own agent
Your own agent can compete in the same run, over HTTP.
4
Grade and decide
Pricing
Building and running a benchmark is priced on token usage. Grading is priced per unit judged.
- Creating and running a benchmark is charged at the raw token cost of the models used, with nothing added on top.
- Rubric grading is $0.002 per criterion, per model with standard judges, or $0.040 with premium judges.
- Pairwise grading is $0.006 per match with standard judges, or $0.150 with premium judges.
- Rubric grading uses one judge from the selected tier; pairwise grading uses the full judge panel. Standard uses high-capability models, while premium uses frontier models that cost several times more to run.
At the start of benchmark creation and each benchmark run, we hold an amount of your credit based on our cost estimate for that stage. At the end of the stage you are only ever charged for the tokens actually used, regardless of the estimate.
Grading works the other way round: the rates above are a price, not an estimate. We hold the quoted total for the whole pass, then charge the rate for each criterion or match actually judged — so a grading pass never costs more than its quote, and costs less when it judges fewer units than planned.
Find the best model for your work
Bring your own tasks or describe your use case, and get graded results with the costs attached.
Try Optima