Build1 distinct publisher3 min readUpdated
Optima lets buyers build benchmarks from their own datasets and agent traces, then scores candidate models on quality, cost per task and time per task.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Artificial Analysis, which built its name on independent LLM evaluations and benchmark implementations such as GDPval-AA and AA-Briefcase [4], has launched Optima, a platform on which customers assemble their own benchmarks from their own data, workflows, or a description of a use case with sample inputs and outputs [1][2]. It is available now [1]. The consequential part is not the custom test set: it is that Optima reports quality, cost per task and time per task as parallel dimensions [3], which is the shape of a bid comparison rather than a leaderboard.
There are three ways in. You can upload existing evaluation datasets from your own files or from Hugging Face, or import agent traces from Arize, Braintrust or Langfuse [5]. Developers can install a skill that collects information from their coding environment and past sessions [5]. If you have none of that, you describe the use case and supply sample inputs and outputs, and Optima proposes test inputs, evaluation criteria and example tasks that you refine through feedback before running anything [6].
Scoring comes in two flavours: a rubric against objective criteria, or the pairwise method Artificial Analysis already uses for GDPval-AA and AA-Briefcase, in which you label a sample of response pairs and the platform derives the full ranking from your stated preferences [7].
The pricing is where the procurement framing gets literal. Artificial Analysis says it passes through the actual token costs of the models used with no markup, and charges $0.125 per criterion per model for rubric evaluations and $0.375 per pairwise comparison [8]. So a 15-criterion rubric across six candidate models is $11.25 in scoring fees on top of token spend [9], and pairwise labelling costs three times as much per unit as a rubric criterion [10]. The platform holds a balance against an estimate at benchmark creation, at each run and at each evaluation round, then bills actual usage [8].
For agentic work, the-decoder argues, headline token price is close to meaningless, because a cheaper model that needs more attempts, fails more often, or leaves cleanup behind can cost more in aggregate, making cost per completed task the number that matters [11]. Artificial Analysis says early testers used the platform to find a model that cut costs for finance and accounting agents by a factor of ten without major quality loss, to identify which model best matched the writing style of lawyers, and to test element identification on a proprietary image dataset [12].
The reason to take this seriously is that public numbers are fragile. An Epoch AI analysis found benchmark results depend on implementation details that are rarely disclosed, with prompt wording and temperature settings alone shifting the same model's score noticeably [13]; on agentic benchmarks such as SWE-bench, swapping the scaffold accounted for up to 15 percentage points [14]. Moving the harness in-house does not eliminate that problem, it relocates it. As the-decoder notes, whether a custom benchmark is methodologically sound and captures actual business value still depends entirely on how it is designed [15].
What to watch: whether procurement teams start writing customer-run cost-per-task thresholds into contracts rather than citing vendor cards; whether pairwise labelling stays affordable once test sets grow past a sample; and whether the per-criterion meter is small enough against the real cost of a wrong model choice, which is migration.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Artificial Analysis has released Optima, a platform that lets users build their own AI benchmarks tailored to specific use cases, and it is available now.
Tests can draw on users' own data sources, workflows, or descriptions of the desired scenario complete with sample inputs and outputs.
Optima compares AI models not just on quality but also on cost per task and time per task, tracked as standalone comparison dimensions.
Artificial Analysis is known for independent LLM evaluations and benchmark implementations including GDPval-AA and AA-Briefcase.
Optima accepts uploaded evaluation datasets from users' own files or from Hugging Face, AI agent traces from platforms such as Arize, Braintrust or Langfuse, and data gathered by an installable skill that collects information from a developer's coding environment and past sessions.
Users without such data can describe their intended use case and provide sample inputs and outputs; Optima then generates suggested test inputs, evaluation criteria and example tasks, which users can review and refine through feedback before running the benchmark.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source launch report, vendor-attributed specifics
Every product and pricing detail traces to one publisher relaying Artificial Analysis's own description ('according to Artificial Analysis'), with no independent test, no reviewer reproduction and no second outlet in the cluster. The pricing and ingestion claims are specific and falsifiable, which lifts the floor, and the article's citation of Epoch AI's configuration-sensitivity findings is external evidence about benchmarking generally rather than about Optima. Nothing verifies that Optima's generated criteria or pairwise extrapolation produce stable rankings.
Available at launch, adoption only vendor anecdote
There is a real availability event -- the platform is shipped and priced, not previewed -- which is more than vaporware. Against that, the only usage signal is unnamed 'early testers' described by the vendor, with no customer names, seat or run counts, revenue, or third-party deployment reports. That supports a low but non-zero measured value.
Mildly overstated, though the article self-corrects
The framing that Optima 'tackles AI benchmarking's biggest flaw' outruns what is shown: a shipped tool with vendor-described early testers and no independent validation, addressing task-fit while leaving reproducibility, criteria validity and business-value measurement untouched. The gap is kept small because the same report explicitly concedes that methodological soundness depends on design, cites Epoch AI on configuration sensitivity and the 445-paper methodology review, and notes cost and time per task do not measure what output is worth.
Vendor-sourced launch for a newly metered product
Artificial Analysis has a direct commercial interest in the claims: Optima is a paid product billing per criterion and per comparison, and the firm's authority derives from the free independent benchmarks it also publishes, so promoting buyer-side custom evaluation both monetises and extends that position. The report is built on the vendor's own description and repeatedly attributes claims to it, and the publisher itself is running a subscription solicitation at the end of the piece. No countervailing sourcing from customers or competitors is present.
Low: one publisher, one vendor voice
Confidence is limited by a single-source, single-publisher cluster in which product, pricing and adoption claims all originate with the vendor. The mechanical details (rates, ingestion paths, scoring modes) are precise enough to be relied on provisionally, and the derived pricing arithmetic is sound, but nothing about efficacy, adoption scale or data handling can be assessed at more than low confidence.
leadership
Re-baseline AI procurement on cost per completed task, not dollars per million tokens1 distinct publisher
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Meta's real announcement is the split: 30B on your GPU, everything else behind the API6 distinct publishers
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.