Skip to content

Build5 publishers3 min readPublished

Anthropic prices Sonnet 5.5 at half of Opus 5.5, with Sonnet close behind on most benchmarks and ahead on Terminal-Bench

Anthropic says Claude Sonnet 5.5 lands within about two points of Opus 5.5 on coding and knowledge-work tests at half Opus's per-token price. The 30 percent saving it advertises is against Sonnet 5, so moving agents off Opus depends on per-task token counts nobody has verified independently.

The Engineer · Build desk

Illustration accompanying Anthropic prices Sonnet 5.5 at half of Opus 5.5, with Sonnet close behind on most benchmarks and ahead on Terminal-Bench

What happened

  • Developers who call Sonnet with thinking turned off must switch to a new between_tools setting before moving to Sonnet 5.5.
  • Sonnet 5.5 is the first Sonnet to ship with Anthropic's top-tier cyber safeguards, and higher-risk security requests fall back to Sonnet 5.
  • Sonnet 5.5 is available now on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Moving routine coding agents from Opus to Sonnet needs a per-task token count from a team's own traces before anyone can say what the switch saves.
  • constraint Effort level has to be chosen per tier and per workload, since the top setting can score below the one beneath it and the low settings may already clear the old model's best.
  • cost Teams that tuned pipelines for thinking-off latency have to re-baseline on both tiers, because Opus 5.5 rejects thinking-off requests and Sonnet 5.5 needs the new between_tools setting.
  • exposure Security tools built on Sonnet will get Sonnet 5 answers on higher-risk requests, so Sonnet 5.5's benchmark gains will not apply evenly to that work.

The 30 percent saving compares Sonnet 5.5 with Sonnet 5 [2]. It comes at an unchanged token price: $2 per million input tokens, $10 per million output tokens and $0.20 per million cache reads [3]. Against Opus 5.5, the difference is in the list price. Opus 5.5 costs $4 and $20 after its cut [4]. Sonnet is half on both lines [1]. On a given task, Sonnet's bill as a fraction of Opus's is half the ratio of their token counts. With the same mix of input and output tokens, Sonnet stays cheaper until it needs twice Opus's tokens for the job [2]. The reports give a cache-read price for Sonnet but not for Opus 5.5 [c3, c4].

Anthropic's reported gaps on the agent-coding tests are small. CursorBench 4.0 recreates real sessions from the Cursor editor. Sonnet 5.5 scored 55.5 percent there against 57.8 for Opus 5.5 [5], a gap of 2.3 points [3]. Cognition's FrontierCode checks whether a code change could merge without human edits. Sonnet scored 52.1 percent at its second-highest effort setting and Opus scored 54.4 [6], again 2.3 points apart [4]. On the GDPval-AA knowledge-work ranking, Sonnet scored 1,844 and Opus 1,846 [7].

The Decoder notes that independent testing still needs to confirm Anthropic's claims [10]. A footnote to the published results says some Sonnet 5.5 values came from a pre-release build with a since-fixed bug that may have affected structured outputs [11]. The Terminal-Bench climb from Sonnet 5 also needs a discount. The benchmark's maintainers said Sonnet 5 sometimes ran into timeouts and token limits [9]. For the CursorBench number to transfer, a team's work has to resemble the Cursor sessions the benchmark recreates [5]. Anthropic draws its own line between the tiers. The company says Sonnet 5.5 is "strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets" [12]. By its account, Opus 5.5 remains "clearly stronger at complex, open-ended work requiring sustained judgment" [13].

Anthropic gives a specific explanation for the Max result. At maximum effort, it says, Sonnet 5.5 more often triggers a code-review function that splits work across sub-agents. Some of those runs timed out or changed code outside the task scope, and FrontierCode penalizes both [15]. I give Anthropic credit for charting its top setting losing to the one below it, with a cause attached. I would expect an agent harness with a wall-clock limit to hit the same failure. At the low end of the effort setting, Anthropic says Sonnet 5.5 at low or medium effort beats Sonnet 5's best score on several benchmarks for about a tenth of the cost per task [17]. It also batches tool calls more often than Sonnet 5, so tasks take fewer steps [16].

What to watch

  • Independent per-task token counts for Sonnet 5.5 against Opus 5.5 on agent workloads; at close to twice Opus's usage, the saving against Opus is gone.
  • Haiku 5.5, due in the coming weeks for high-throughput, low-cost work, adds a third tier to the same routing test.
  • Whether Anthropic reissues the Sonnet 5.5 figures from the release build after the since-fixed structured-output bug.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories