BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Failed runs decide whether Claude Sonnet 5.5 undercuts Opus 5.5
Anthropic prices Claude Sonnet 5.5 at half Opus 5.5's per-token rate, but Artificial Analysis measured $7.67 per task at max effort against Opus's $5.98. In The New Stack's own tests, runs that were billed but never finished decided which model cost less per completed job.
The Engineer · Build desk

What happened
- Sonnet 5.5 overran a 32,000-token per-step limit in four of five runs of The New Stack's agentic bug-fix test and passed only after the cap rose to 128,000.
- On a concurrency-bug task with eight hidden tests, Sonnet passed every check in all five runs, while Opus 5.5 matched it in only three.
- On Terminal-Bench 4.0, Anthropic puts Sonnet 5.5 at 70.6% and Opus 5.5, run at xhigh effort, at 66.4%.
Why it matters
- cost Budgets built from list price undercount agentic spend, because a run cut off at the step cap is still billed and produces nothing usable.
- decision Picking a model now means picking its per-step output cap as well, since the limit that truncated Sonnet was one Opus never approached.
- contradiction Artificial Analysis and The New Stack disagree on which model is cheaper per task, so neither result can stand in for a team's own measurement on its own jobs.
Per-task cost is token price multiplied by tokens consumed. Sonnet 5.5 lists at exactly half of Opus 5.5 on both input and output [13], so it breaks even when it spends twice the tokens [21]. The New Stack ran three tests through the Anthropic API at maximum effort with adaptive thinking. Each model got five runs per test, graded against hidden tests the models never saw [4]. On the resolver spec, Sonnet used 15% more output tokens than Opus and came in 42% cheaper [18]. On the agentic bug fix it used 2.07 times as many [14], and its saving fell to about 7% a run before any failed runs were counted [15].
The failed runs trace to one harness setting. Opus ran under the same 32,000-token step cap and never came close to it [7]. Both models ended up fixing all four planted bugs on every run [5]. Sonnet's truncated first attempts cost about $1.40, and The New Stack left them out of its reported averages [8]. Spread across five runs, that adds $0.28 each and lifts Sonnet from $0.70 to $0.98 against Opus's $0.75 [16]. A 7% saving becomes a 31% premium [22].
A higher cap has its own failure. Twice on the concurrency task, Opus spent 128,000 tokens deliberating over three race conditions and kept its conclusions to itself [10]. Those empty runs are inside its $2.24 per-run average [11]. Five runs at $2.24 is $11.20, and only three produced a passing fix, so each completed job cost about $3.73 [17]. Sonnet's $1.02 a run bought a passing fix every time [c12, c13].
Speed split as well. Opus finished the bug fix in about 35% less time [20]. Sonnet was faster on the resolver and on the concurrency task [c11, c13].
The two independent measurements disagree. Artificial Analysis's max-effort figures put Sonnet 28% above Opus per task [12]. The New Stack's concurrency runs had Sonnet about 54% cheaper per run [19]. The New Stack's article does not describe Artificial Analysis's task set, so the record cannot say which workload each result resembles. The New Stack's own sample is 15 runs per model across three tests [23]. For any of these figures to carry over, a team's jobs would need to match the setup: maximum effort, a per-step cap near 32,000 or 128,000 tokens, and pass-or-fail grading against hidden tests [c6, c9].
In my view the comparison belongs at the level of cost per completed job, measured on a sample of a team's own tasks with failed runs left in the total. By that measure The New Stack's tests went to Sonnet twice and to Opus once [d11, d7, d6].
What to watch
- Artificial Analysis publishing the task mix behind its $7.67 and $5.98 per-task figures, to show whether its workload resembles multi-step agentic work.
- Per-task costs for both models at effort settings below maximum, since every figure here was measured at max effort.
- Whether Opus 5.5 completes the two concurrency runs it lost when given an output budget above 128,000 tokens.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives35
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Independent testing from Artificial Analysis found that at max effort Sonnet 5.5 cost $7.67 per task to Opus 5.5's $5.98.
ReportedSupportedSource: Artificial Analysis, as reported by The New Stack2 sources— create a free account to open themView cited source - [2]
Anthropic says Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5's 66.4% at xhigh effort.
- [3]
Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens; Opus 5.5 costs $4 and $20.
- [4]
The New Stack ran three tests, calling both models through the Anthropic API with identical prompts, adaptive thinking and the maximum effort setting, five runs per model per test, grading every run against a hidden test suite the models never saw and logging tokens, list-price cost and time.
- [5]
In the agentic bug-fix test, both models fixed all four planted bugs in all five runs and passed all 12 hidden tests.
- [6]
In the agentic bug-fix test, Sonnet 5.5 averaged 5 minutes 8 seconds, 29 tool calls, 42,608 output tokens and $0.70 per run; Opus 5.5 averaged 3 minutes 21 seconds, 25 tool calls, 20,625 output tokens and $0.75.
- [7]
In the agentic test each step had a 32,000-token limit; on the first attempt Sonnet 5.5 thought so long in a single step that it hit the limit in four of five runs, which stopped before finishing. Opus 5.5 ran under the same limit and never came close. The limit was raised to 128,000, the four were rerun, and all passed.
- [8]
Sonnet 5.5's failed agentic attempts cost about $1.40, not included in the totals; counting them, Sonnet 5.5 averaged about $0.98 per run on the test, more than Opus 5.5's $0.75.
- [9]
On the resolver spec test, both models passed all 120 hidden tests on all five runs; Sonnet 5.5 averaged 8 minutes 58 seconds, 81,097 output tokens and $0.82 per run, Opus 5.5 averaged 9 minutes 40 seconds, 70,687 output tokens and $1.42.
- [10]
On the concurrency test, Sonnet 5.5 fixed all three race conditions and passed all eight hidden tests on all five runs; Opus 5.5 matched it on three runs and on the other two spent all 128,000 output tokens thinking and never produced an answer.
- [11]
On the concurrency test, Sonnet 5.5 averaged 12 minutes 13 seconds, 101,788 output tokens and $1.02 per run; Opus 5.5 averaged 16 minutes 43 seconds, 111,428 output tokens and $2.24 per run including failures.
- [12]
Artificial Analysis's figures put Sonnet 5.5 about 28% above Opus 5.5 per task at max effort.
- [13]
Sonnet 5.5's list price is 50% of Opus 5.5's on both input and output tokens.
- [14]
On the agentic bug fix, Sonnet 5.5 used 2.07 times as many output tokens as Opus 5.5.
- [15]
Excluding failed attempts, Sonnet 5.5 was about 7% cheaper per run than Opus 5.5 on the agentic bug fix.
- [16]
Spreading the $1.40 of failed attempts across five runs adds $0.28 per run, taking Sonnet 5.5 from $0.70 to $0.98.
- [17]
Opus 5.5 cost about $3.73 per completed concurrency fix: $11.20 across five runs divided by three passing runs.
- [18]
On the resolver spec, Sonnet 5.5 used about 15% more output tokens than Opus 5.5 and cost about 42% less per run.
- [19]
On the concurrency test, Sonnet 5.5 cost about 54% less per run than Opus 5.5.
- [20]
Opus 5.5 finished the agentic bug fix in about 35% less time than Sonnet 5.5.
- [21]
At half the per-token price, Sonnet 5.5 breaks even with Opus 5.5 when it spends twice the tokens.
- [22]
Counting failed attempts, Sonnet 5.5 cost about 31% more per run than Opus 5.5 on the agentic bug fix.
- [23]
The New Stack's sample is 15 runs per model.
Sources
1 independent publisher whose own reporting we read for this story.
- thenewstack.ioClaude Sonnet 5.5 vs. Opus 5.5: 42% cheaper and perfect on every run
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.