Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Failed runs decide whether Claude Sonnet 5.5 undercuts Opus 5.5

Anthropic prices Claude Sonnet 5.5 at half Opus 5.5's per-token rate, but Artificial Analysis measured $7.67 per task at max effort against Opus's $5.98. In The New Stack's own tests, runs that were billed but never finished decided which model cost less per completed job.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Failed runs decide whether Claude Sonnet 5.5 undercuts Opus 5.5
Generated illustration

What happened

  • Sonnet 5.5 overran a 32,000-token per-step limit in four of five runs of The New Stack's agentic bug-fix test and passed only after the cap rose to 128,000.
  • On a concurrency-bug task with eight hidden tests, Sonnet passed every check in all five runs, while Opus 5.5 matched it in only three.
  • On Terminal-Bench 4.0, Anthropic puts Sonnet 5.5 at 70.6% and Opus 5.5, run at xhigh effort, at 66.4%.

Why it matters

  • cost Budgets built from list price undercount agentic spend, because a run cut off at the step cap is still billed and produces nothing usable.
  • decision Picking a model now means picking its per-step output cap as well, since the limit that truncated Sonnet was one Opus never approached.
  • contradiction Artificial Analysis and The New Stack disagree on which model is cheaper per task, so neither result can stand in for a team's own measurement on its own jobs.

Per-task cost is token price multiplied by tokens consumed. Sonnet 5.5 lists at exactly half of Opus 5.5 on both input and output [13], so it breaks even when it spends twice the tokens [21]. The New Stack ran three tests through the Anthropic API at maximum effort with adaptive thinking. Each model got five runs per test, graded against hidden tests the models never saw [4]. On the resolver spec, Sonnet used 15% more output tokens than Opus and came in 42% cheaper [18]. On the agentic bug fix it used 2.07 times as many [14], and its saving fell to about 7% a run before any failed runs were counted [15].

The failed runs trace to one harness setting. Opus ran under the same 32,000-token step cap and never came close to it [7]. Both models ended up fixing all four planted bugs on every run [5]. Sonnet's truncated first attempts cost about $1.40, and The New Stack left them out of its reported averages [8]. Spread across five runs, that adds $0.28 each and lifts Sonnet from $0.70 to $0.98 against Opus's $0.75 [16]. A 7% saving becomes a 31% premium [22].

A higher cap has its own failure. Twice on the concurrency task, Opus spent 128,000 tokens deliberating over three race conditions and kept its conclusions to itself [10]. Those empty runs are inside its $2.24 per-run average [11]. Five runs at $2.24 is $11.20, and only three produced a passing fix, so each completed job cost about $3.73 [17]. Sonnet's $1.02 a run bought a passing fix every time [c12, c13].

Speed split as well. Opus finished the bug fix in about 35% less time [20]. Sonnet was faster on the resolver and on the concurrency task [c11, c13].

The two independent measurements disagree. Artificial Analysis's max-effort figures put Sonnet 28% above Opus per task [12]. The New Stack's concurrency runs had Sonnet about 54% cheaper per run [19]. The New Stack's article does not describe Artificial Analysis's task set, so the record cannot say which workload each result resembles. The New Stack's own sample is 15 runs per model across three tests [23]. For any of these figures to carry over, a team's jobs would need to match the setup: maximum effort, a per-step cap near 32,000 or 128,000 tokens, and pass-or-fail grading against hidden tests [c6, c9].

In my view the comparison belongs at the level of cost per completed job, measured on a sample of a team's own tasks with failed runs left in the total. By that measure The New Stack's tests went to Sonnet twice and to Opus once [d11, d7, d6].

What to watch

  • Artificial Analysis publishing the task mix behind its $7.67 and $5.98 per-task figures, to show whether its workload resembles multi-step agentic work.
  • Per-task costs for both models at effort settings below maximum, since every figure here was measured at max effort.
  • Whether Opus 5.5 completes the two concurrency runs it lost when given an output budget above 128,000 tokens.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption
Insufficient
Hype gap+20
Incentives35
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Independent testing from Artificial Analysis found that at max effort Sonnet 5.5 cost $7.67 per task to Opus 5.5's $5.98.

    ReportedSupportedSource: Artificial Analysis, as reported by The New Stack2 sources— create a free account to open themView cited source
  2. [2]

    Anthropic says Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5's 66.4% at xhigh effort.

    ReportedSupportedSource: Anthropic, as reported by The New StackView cited source
  3. [3]

    Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens; Opus 5.5 costs $4 and $20.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. thenewstack.io

    1 article · October 8, 2026

    Claude Sonnet 5.5 vs. Opus 5.5: 42% cheaper and perfect on every run

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories