Leadership1 distinct publisher3 min readUpdated
Arize and Fireworks ran ten models against 40 agent tasks and found the cheapest model per finished job also had the worst pass rate. Coverage, not price, is the binding constraint.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
Arize and Fireworks have published a joint benchmark of ten open and closed models from four providers, run against 40 real command-line agent tasks at six trials each for 2,400 runs [1][2]. Their headline finding is a procurement claim rather than a capability one: the open-versus-closed and cheap-versus-frontier labels teams use to shortlist models did not predict which model finished work for less money, and what separated them was cost per successful task [3].
The metric is defined as everything a model spent across every attempt, including runs that failed, retried, or timed out, divided by the number of times it actually completed the task [4]. The point of the definition is what a pricing page cannot see: retries, failed tool calls, malformed outputs, judge calls, and runs that grind to a limit and never finish [5]. Arize calls the attempts-needed-per-success figure the retry tax [6].
The result that carries the argument is also the most awkward one. According to Arize, gpt-oss-120b posted the worst pass rate in the study at 33 percent and still finished a task for $0.054, roughly 12 times cheaper per success than GPT-5.5 and 23 times cheaper than gemini-3.5-flash [7]. Taken at face value, that puts the Gemini model near $1.24 per completed task [8].
Splitting by difficulty is where this becomes usable. On easy tasks the frontier premium buys nothing: GPT-5.5 passed 69 percent, Kimi K2.6 passed 73 percent, and gpt-oss-120b passed 65 percent at roughly a twenty-fourth of the cost per attempt [9]. On hard tasks the cheap models fall away and only frontier-class models compete, open or closed [10]. Arize is blunt about the limit of retry-based savings: "a model that cannot do a task does not learn it on the fourth attempt, it just bills you four times" [11].
Frontier-class is not one thing either. GPT-5.5 and Kimi K3 finished in a near-tie at 67 and 66 percent success and $0.64 and $0.67 per successful task, but with opposite strengths [12]. K3 solved every easy task in all six trials while GPT-5.5 slipped to 69 percent; on the hardest tasks GPT-5.5 hit 51 percent against K3's 32 [13][14], a 19-point swing in the other direction [15].
The counterweight to the cheap-model result is coverage. Counting tasks solved reliably, meaning four of six trials or better, gpt-oss-120b managed 8 of 40 against 25 for GPT-5.5 and 26 for Kimi K3 [16], or 20 percent of the task set versus 63 and 65 percent [17]. Cost per success flatters a model that only enters races it can win; swap it in wholesale, Arize argues, and the bill collapses along with the set of things the product can do [18].
The framing matters because Arize expects the subsidised-inference era to end, at which point the gap between a model that looks cheap on a pricing page and one that is cheap to finish work with becomes the difference between a product that pencils out and one that does not [19].
Caveats worth holding: this is a single benchmark self-published by the two firms that ran it, Kimi K3 is not yet open-weights and was tested through Kimi.com [20], and Arize explicitly excludes GPT-5.6 Sol from the comparison [21]. Watch whether K3's weights ship as promised, whether any vendor starts quoting attempts per success alongside price per million tokens, and whether routing products publish a coverage floor rather than a savings percentage.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Arize and Fireworks tested 10 open and closed models from four providers on 40 real command-line agent tasks, measuring pass rates, retries, coverage, and cost per successful task.
The study comprised 2,400 runs: 40 tasks, 10 models, 6 trials each.
gpt-oss-120b, an open model, was the most cost effective per successful task and had the worst pass rate in the study: it finished a task for $0.054 while passing only 33% of the time, roughly 12x lower cost per success than GPT-5.5 and 23x more cost effective than gemini-3.5-flash.
On hard tasks the most cost-effective models fall far behind and only frontier-class models, open or closed, compete; Arize says this is the one place routing cannot save money.
Arize writes: "a model that cannot do a task does not learn it on the fourth attempt, it just bills you four times."
Cost per successful task is defined as everything a model spent across every attempt, including runs that failed, retried, or timed out, divided by the number of times it completed the task.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific methodology, single self-reported run
The benchmark is described with unusual specificity — 2,400 runs across 40 Terminal-Bench tasks with per-task Docker environments and test scripts, a deliberately thin one-tool agent, one code path per provider, and an explicit cost-per-success definition — and the authors caveat that GPT-5.6 Sol was excluded. But everything is self-reported by the two vendors who ran it: the supplied text contains no raw runs, no per-task variance or confidence intervals, no published price table behind the dollar figures, and no independent replication. Evidence is therefore credible on design and thin on verification.
Confined to the authoring vendors
The only observed usage is the benchmark the authors ran themselves: ten models exercised through their own harness, including Kimi K3 accessed via Kimi.com. There is no evidence in the supplied source of any third party adopting cost per successful task as a selection metric, of routing being deployed in production on the basis of these results, or of other teams reproducing the runs.
Framing outruns a single unreplicated run
The interpretive layer — that open/closed labels are 'the wrong axis', that the shortlist should be reordered, and that the subsidised-inference era is ending — is broader than one vendor-run benchmark on 40 command-line tasks can carry, and the subsidy claim arrives with no pricing or margin evidence at all. The gap is moderate rather than large because the post volunteers the strongest counterweight to its own cheap-model headline (8 of 40 reliable coverage), states the hard-task limit of routing, and flags the excluded model.
Vendor-authored on a thesis both sellers benefit from
The study is authored by Arize and Fireworks and published on Arize's blog. The conclusions — measure cost per finished job, instrument retries, and route between cheap and frontier models — align with the measurement and serving businesses of the two authors, and the finding that open models occupy the efficient frontier is favourable to a provider that serves open-weight models such as gpt-oss-120b. The post does not disclose these incentives, and no independent party in the cluster checks the numbers.
Single publisher, precise but unverified
Confidence is limited by cluster structure: one publisher, one document, no counter-coverage and no reproduction. The numbers are precise and internally consistent, and the methodology is described well enough to be re-run, which keeps confidence from falling further; but the results are vendor-produced on a self-selected model list in one task domain, so the durable finding is the metric rather than the ranking.
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
build
PerceptionBench puts a number on the step your pipeline treats as free1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026