Build1 publisher3 min readPublished
Artificial Analysis moves eval onto your data, and turns model choice into procurement
Optima lets buyers build benchmarks from their own datasets and agent traces, then scores candidate models on quality, cost per task and time per task.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Artificial Analysis has released Optima, a platform that lets users build their own AI benchmarks tailored to specific use cases, and it is available now.
- Tests can draw on users' own data sources, workflows, or descriptions of the desired scenario complete with sample inputs and outputs.
- Optima compares AI models not just on quality but also on cost per task and time per task, tracked as standalone comparison dimensions.
- Artificial Analysis is known for independent LLM evaluations and benchmark implementations including GDPval-AA and AA-Briefcase.
- Optima accepts uploaded evaluation datasets from users' own files or from Hugging Face, AI agent traces from platforms such as Arize, Braintrust or Langfuse, and data gathered by an installable skill that collects information from a developer's coding environment and past sessions.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Artificial Analysis, which built its name on independent LLM evaluations and benchmark implementations such as GDPval-AA and AA-Briefcase [4], has launched Optima, a platform on which customers assemble their own benchmarks from their own data, workflows, or a description of a use case with sample inputs and outputs [1][2]. It is available now [1]. The consequential part is not the custom test set: it is that Optima reports quality, cost per task and time per task as parallel dimensions [3], which is the shape of a bid comparison rather than a leaderboard.
There are three ways in. You can upload existing evaluation datasets from your own files or from Hugging Face, or import agent traces from Arize, Braintrust or Langfuse [5]. Developers can install a skill that collects information from their coding environment and past sessions [5]. If you have none of that, you describe the use case and supply sample inputs and outputs, and Optima proposes test inputs, evaluation criteria and example tasks that you refine through feedback before running anything [6].
Scoring comes in two flavours: a rubric against objective criteria, or the pairwise method Artificial Analysis already uses for GDPval-AA and AA-Briefcase, in which you label a sample of response pairs and the platform derives the full ranking from your stated preferences [7].
The pricing is where the procurement framing gets literal. Artificial Analysis says it passes through the actual token costs of the models used with no markup, and charges $0.125 per criterion per model for rubric evaluations and $0.375 per pairwise comparison [8]. So a 15-criterion rubric across six candidate models is $11.25 in scoring fees on top of token spend [9], and pairwise labelling costs three times as much per unit as a rubric criterion [10]. The platform holds a balance against an estimate at benchmark creation, at each run and at each evaluation round, then bills actual usage [8].
For agentic work, the-decoder argues, headline token price is close to meaningless, because a cheaper model that needs more attempts, fails more often, or leaves cleanup behind can cost more in aggregate, making cost per completed task the number that matters [11]. Artificial Analysis says early testers used the platform to find a model that cut costs for finance and accounting agents by a factor of ten without major quality loss, to identify which model best matched the writing style of lawyers, and to test element identification on a proprietary image dataset [12].
The reason to take this seriously is that public numbers are fragile. An Epoch AI analysis found benchmark results depend on implementation details that are rarely disclosed, with prompt wording and temperature settings alone shifting the same model's score noticeably [13]; on agentic benchmarks such as SWE-bench, swapping the scaffold accounted for up to 15 percentage points [14]. Moving the harness in-house does not eliminate that problem, it relocates it. As the-decoder notes, whether a custom benchmark is methodologically sound and captures actual business value still depends entirely on how it is designed [15].
What to watch: whether procurement teams start writing customer-run cost-per-task thresholds into contracts rather than citing vendor cards; whether pairwise labelling stays affordable once test sets grow past a sample; and whether the per-criterion meter is small enough against the real cost of a wrong model choice, which is migration.