Skip to content

Invest2 publishers3 min readPublished

Sonnet 5.5's extra tokens shrink its half-price edge over Opus to roughly a fifth at max effort

Anthropic's Claude Sonnet 5.5 beats Opus 5.5 at coding for half the per-token price, on its own tests and on Artificial Analysis's. At max effort it writes 60% more tokens per task, so moving coding work down a tier saves nearer a fifth than a half.

The Investor · Invest desk

Illustration accompanying Sonnet 5.5's extra tokens shrink its half-price edge over Opus to roughly a fifth at max effort

What happened

  • Anthropic says that at Medium effort, the default in its apps, Sonnet 5.5 beats Sonnet 5's best coding score for less than a tenth of the cost.
  • Sonnet 5.5's best CursorBench score came within about two points of Opus 5.5, according to Anthropic.
  • Artificial Analysis tested a pre-release build with a bug that Anthropic expects changed the scores little or slightly understated them.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • decision Bug fixes and other scoped coding tickets can move from Opus to Sonnet at High effort, the setting Artificial Analysis rates best value, while Opus budgets go to open-ended judgment work.
  • cost Because Sonnet 5.5's per-token price is unchanged from Sonnet 5, every saving Anthropic claims comes from using fewer tokens, and a workload that pushes effort up spends that saving.
  • constraint With OpenAI's flagship now at Sonnet's per-token price, list prices stop separating the two vendors; buyers have to compare tokens per task, where Sonnet 5.5 is the heaviest model Artificial Analysis has measured.

The independent coding win and the heaviest bill come from the same setting. Artificial Analysis's Terminal-Bench score is for Sonnet 5.5 at max effort [4]. At max effort the model wrote about 193,000 tokens per test task, the most the firm has measured and roughly 60% more than Opus 5.5 [6]. Half the per-token price times 1.6 times the tokens is 0.8. On that basis a max-effort Sonnet task costs about 20% less than the same task on Opus [2]. Opus would run near $9.50 a task against Sonnet's $7.60 [3].

That estimate assumes the token gap runs through the whole bill, and the figures leave a puzzle. At the $10 output rate [1], 193,000 written tokens come to about $1.93, roughly a quarter of $7.60 [4]. The rest is presumably input the model reads back on each step at $2 per million, though the source does not break the bill down. If Sonnet's longer runs also re-read more context, the saving stays near a fifth. If its input volume matches Opus's, the saving widens, since each input token costs half as much.

Both published savings figures can hold at once. Anthropic says Sonnet 5.5 costs up to 30% less per task than Sonnet 5 because it needs fewer tokens [8], while Artificial Analysis's max-effort bill points the other way [7]. The difference is the effort setting. Artificial Analysis calls High effort the best value [10], and Anthropic says that at High, Sonnet 5.5 matches GPT-6 Sol on FrontierCode for about a fifth of the cost per task [11].

The coding lead holds up under outside testing. Sonnet's margin over Opus on Terminal-Bench 4.0 is 4.2 points on Anthropic's self-reported table [20] and 4.0 points on Artificial Analysis's run [5], even though the independent scores sit 7 points lower [6]. Overall, the firm ranks Sonnet second behind Opus [5]. "Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It's also got a sharp eye for design," Anthropic wrote [13]. The company says Opus 5.5 remains clearly stronger at complex work needing sustained judgment [14].

Anthropic held Sonnet's list price flat from Sonnet 5 [1] and left Opus at twice that, about $4 and $20 per million tokens [1]. It published no benchmarks against GPT-5.6 Terra, OpenAI's mid-tier model [16].

The result turns on where a team sets the dial. It can run Sonnet at Medium or High and keep most of the price gap. It can find it needs max effort to match Opus on its own tickets and save about a fifth. Or a retest could revise Artificial Analysis's figures, which came from a pre-release build with a bug [17]. In my view the first case fits bug fixes and scoped tickets, with Opus kept for the open-ended work Anthropic itself assigns it. The counter-case is that real tickets are messier than benchmark tasks, and teams drift to max effort anyway. The view is wrong if a team's own logs show Sonnet reaching Opus-level pass rates only at max effort, with a per-task bill within a fifth of Opus's.

What to watch

  • An Artificial Analysis retest on the release build: a max-effort token count below 193,000 per task would widen Sonnet's per-task saving over Opus.
  • Pricing and token use for Claude Haiku 5.5, which Anthropic says is due in the coming weeks for high-volume, cost-sensitive work.
  • Head-to-head coding results against GPT-5.6 Terra, which lists at $2 and $12 and which Anthropic left out of its benchmarks.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories