Published · 5d agoLeadership2 min read
The $0.054 Task: The Cheapest Model Per Success Covers A Fifth Of The Work
Arize and Fireworks scored ten models on cost per successful task across 2,400 agent runs. The cheapest finisher also had the worst pass rate, which makes the number real and the substitution unwise.
Context for builders, not their beat.See today for builders

What happened
- Arize and Fireworks tested 10 open and closed models on 40 real command-line agent tasks across 2,400 runs (40 tasks, 10 models, 6 trials each), measuring pass rates, retries, coverage and cost per successful task rather than token price.
- gpt-oss-120b finishes a task for $0.054, roughly 12x lower cost per success than GPT-5.5 and 23x more cost effective than gemini-3.5-flash.
- Cost per successful task is defined as everything a model spent across every attempt, including runs that failed, retried or timed out, divided by the number of times it actually completed the task.
- Token price cannot see retries, failed tool calls, malformed outputs, judge calls, or runs that grind to a limit and never finish; cost per successful task sees all of it.
- The most cost effective model per successful task in the study is an open one and has the worst pass rate in the study, passing only 33% of the time.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
Arize and Fireworks ran ten open and closed models through 40 real command-line agent tasks, 2,400 runs in total, and ranked them on cost per successful task instead of cost per token [1]. The cheapest finisher was gpt-oss-120b at $0.054 per completed task, roughly 12 times lower than GPT-5.5 and 23 times more cost effective than gemini-3.5-flash [2].
The metric earns its keep because it swallows the failures: everything spent across every attempt, including runs that failed, retried or timed out, divided by the completions [3]. Token price cannot see retries, malformed outputs or runs that grind to a limit [4]. What cost per success still cannot see is coverage. gpt-oss-120b also had the worst pass rate in the study at 33% [5], and solved just 8 of the 40 tasks reliably, meaning 4 of 6 trials or better, against 25 for GPT-5.5 and 26 for Kimi K3 [6]. That is a fifth of the task set versus roughly two thirds [7]. Arize's conclusion is that swapping wholesale collapses your bill and the set of things your product can do [8].
The split by difficulty is where the procurement decision actually sits. On easy work the frontier premium buys nothing: GPT-5.5 passes 69%, Kimi K2.6 73%, and gpt-oss-120b 65% at roughly a twenty-fourth of the cost per attempt [9]. On hard work only the top tier competes, and retrying is not a substitute, because a model that cannot do a task just bills you four times for the attempt [10]. Even within that tier the labels blur: GPT-5.5 and Kimi K3 finished in a near-tie at 67% and 66% success for $0.64 and $0.67 per successful task, but K3 solved every easy task while GPT-5.5 led on the hardest, 51% to 32% [11].
Buyers are already reading this way. OpenAI CFO Sarah Friar said enterprise sales passed consumer revenue, taking annualized revenue to $40 billion after a 20 percent July rise, with business accounts up 32 percent to two million and corporate buyers focused on real work per dollar rather than raw token counts [12]. Exponential View's tracking points the same direction: top-model usage inside businesses has gone flat at 6% of tokens and 11% of spend, which it reads as a ceiling on willingness to pay for the best model [13].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Arize and Fireworks tested 10 open and closed models on 40 real command-line agent tasks across 2,400 runs (40 tasks, 10 models, 6 trials each), measuring pass rates, retries, coverage and cost per successful task rather than token price.
ReportedView cited source - [2]
gpt-oss-120b finishes a task for $0.054, roughly 12x lower cost per success than GPT-5.5 and 23x more cost effective than gemini-3.5-flash.
ReportedView cited source - [3]
Cost per successful task is defined as everything a model spent across every attempt, including runs that failed, retried or timed out, divided by the number of times it actually completed the task.
ReportedView cited source - [4]
Token price cannot see retries, failed tool calls, malformed outputs, judge calls, or runs that grind to a limit and never finish; cost per successful task sees all of it.
ReportedView cited source - [5]
The most cost effective model per successful task in the study is an open one and has the worst pass rate in the study, passing only 33% of the time.
ReportedView cited source - [6]
Counting tasks a model solves reliably (4 of 6 trials or better): gpt-oss-120b 8 of 40, GPT-5.5 25 of 40, Kimi K3 26 of 40.
ReportedView cited source
Sources & coverage · 6 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- aibreakfast.beehiiv.com5d agoOpenAI adds opt-in desktop activity logging for context
- exponentialview.co5d ago📈 Data to start your week


