Leadership1 publisher3 min readPublished
Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist
Arize and Fireworks ran ten models against 40 agent tasks and found the cheapest model per finished job also had the worst pass rate. Coverage, not price, is the binding constraint.
The Board Room · Leadership desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Arize and Fireworks tested 10 open and closed models from four providers on 40 real command-line agent tasks, measuring pass rates, retries, coverage, and cost per successful task.
- The study comprised 2,400 runs: 40 tasks, 10 models, 6 trials each.
- Arize and Fireworks argue the open/closed and cost-effective/frontier labels turned out to be the wrong axis, and that what separated the models was cost per successful task.
- Cost per successful task is defined as everything a model spent across every attempt, including runs that failed, retried, or timed out, divided by the number of times it completed the task.
- Token price cannot see retries, failed tool calls, malformed outputs, judge calls, or runs that grind to a limit and never finish; cost per successful task captures all of it.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
Arize and Fireworks have published a joint benchmark of ten open and closed models from four providers, run against 40 real command-line agent tasks at six trials each for 2,400 runs [1][2]. Their headline finding is a procurement claim rather than a capability one: the open-versus-closed and cheap-versus-frontier labels teams use to shortlist models did not predict which model finished work for less money, and what separated them was cost per successful task [3].
The metric is defined as everything a model spent across every attempt, including runs that failed, retried, or timed out, divided by the number of times it actually completed the task [4]. The point of the definition is what a pricing page cannot see: retries, failed tool calls, malformed outputs, judge calls, and runs that grind to a limit and never finish [5]. Arize calls the attempts-needed-per-success figure the retry tax [6].
The result that carries the argument is also the most awkward one. According to Arize, gpt-oss-120b posted the worst pass rate in the study at 33 percent and still finished a task for $0.054, roughly 12 times cheaper per success than GPT-5.5 and 23 times cheaper than gemini-3.5-flash [7]. Taken at face value, that puts the Gemini model near $1.24 per completed task [8].
Splitting by difficulty is where this becomes usable. On easy tasks the frontier premium buys nothing: GPT-5.5 passed 69 percent, Kimi K2.6 passed 73 percent, and gpt-oss-120b passed 65 percent at roughly a twenty-fourth of the cost per attempt [9]. On hard tasks the cheap models fall away and only frontier-class models compete, open or closed [10]. Arize is blunt about the limit of retry-based savings: "a model that cannot do a task does not learn it on the fourth attempt, it just bills you four times" [11].
Frontier-class is not one thing either. GPT-5.5 and Kimi K3 finished in a near-tie at 67 and 66 percent success and $0.64 and $0.67 per successful task, but with opposite strengths [12]. K3 solved every easy task in all six trials while GPT-5.5 slipped to 69 percent; on the hardest tasks GPT-5.5 hit 51 percent against K3's 32 [13][14], a 19-point swing in the other direction [15].
The counterweight to the cheap-model result is coverage. Counting tasks solved reliably, meaning four of six trials or better, gpt-oss-120b managed 8 of 40 against 25 for GPT-5.5 and 26 for Kimi K3 [16], or 20 percent of the task set versus 63 and 65 percent [17]. Cost per success flatters a model that only enters races it can win; swap it in wholesale, Arize argues, and the bill collapses along with the set of things the product can do [18].
The framing matters because Arize expects the subsidised-inference era to end, at which point the gap between a model that looks cheap on a pricing page and one that is cheap to finish work with becomes the difference between a product that pencils out and one that does not [19].
Caveats worth holding: this is a single benchmark self-published by the two firms that ran it, Kimi K3 is not yet open-weights and was tested through Kimi.com [20], and Arize explicitly excludes GPT-5.6 Sol from the comparison [21]. Watch whether K3's weights ship as promised, whether any vendor starts quoting attempts per success alongside price per million tokens, and whether routing products publish a coverage floor rather than a savings percentage.