Invest4 publishers3 min readPublished Updated
Cached context shrinks the discount from Anthropic's half-price Sonnet 5.5
Anthropic released Sonnet 5.5 at $2 and $10 per million input and output tokens, half the Opus 5.5 rate. How much a buyer saves by moving work down a tier depends on tokens burned per task and on cache reads priced identically on both models.
The Investor · Invest desk

What happened
- Sonnet 5.5 is the second model in the Claude 5.5 family, launched less than a week after Opus 5.5.
- Haiku 5.5, the cheapest tier and built for high-volume, cost-sensitive use, is due to join the family in the coming weeks.
- Sonnet 5.5 scored 70.6% on the Terminal-Bench 4.0 coding test, against 10.3% for Sonnet 5.
- Anthropic called the model's cyber capabilities a "large improvement" and shipped it as the first Sonnet with safeguards built for its most capable models.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- decision Teams splitting jobs between Sonnet and Opus have to count tokens per task on their own workloads before the half-price rate says anything about their bill.
- constraint Agents that reread long cached contexts get the smallest discount from moving down a tier, because the cached share of the bill costs the same on either model.
- cost Customers already on Sonnet 5 keep paying the same rate and capture savings only as jobs shrink in tokens, so Anthropic's per-token revenue on the tier stays intact.
- exposure Routine Opus 5.5 volume is the revenue most exposed to the split, since Anthropic's own product staff pitch Sonnet to customers who may not need Opus's judgment.
Per token, Sonnet 5.5 costs half of Opus 5.5 on input and output alike: $2 against $4 per million input tokens and $10 against $20 per million output [1][2]. Per task, the comparison depends on token count. Sonnet comes out cheaper on a given job only while it burns fewer than twice the tokens Opus would on that job [1]. Anthropic's efficiency claim is a cost per task up to 30% lower at an unchanged per-token price, and it is measured against Sonnet 5 [5]. The sources do not include a same-task token count against Opus 5.5, and a buyer splitting work between the two needs exactly that ratio.
Cache reads narrow the gap. Both models charge $0.20 per million cached tokens [2]. Take an agent job that rereads 10 million cached tokens and writes 100,000 output tokens, leaving fresh input aside. Opus bills $2.00 for the cache and $2.00 for the output, Sonnet bills $2.00 and $1.00, and moving the job down a tier saves 25% [2].
Anthropic draws the line between tiers at judgment. It says Opus 5.5 is meant for complex work that needs careful judgment, while Sonnet 5.5 is strongest at well-scoped everyday tasks, bug fixes and polished documents, slides and spreadsheets [14]. "Sonnet is really for the cost-conscious customer where they might not need as much intelligence," Theo Chu, a research product manager at Anthropic, told CNBC [10]. "It might be routine tasks that just need execution, but don't need that judgment that Opus can bring," he said [11]. On GDPval-AA, a test of real-world work across occupations, the model with double the token price scored two points higher [7][1].
Anthropic did not cut Sonnet's rate to get there. The $2 and $10 prices were Sonnet 5's too [1], so a job that cost $1.00 on Sonnet 5 costs as little as $0.70 on Sonnet 5.5, and all of that saving comes from fewer tokens [4]. Early testers reported gains in the same direction. Box said the model was 2.4 times faster and used 12% fewer tokens on sensitive document work [8], and Base44 said it needed under half the iterations Sonnet 5 required to finish real app builds [9].
The tiering reaches Anthropic's safety spending as well. Because Sonnet 5.5 does not advance the frontier of its models' capabilities, the company said, most of its alignment testing focused on a "targeted set of risks that apply to models of any capability level" [12].
The split can land on Anthropic's revenue in more than one way. Opus customers could move routine jobs to Sonnet and pay half per token on them. Sonnet 5.5 could instead mostly absorb Sonnet 5 traffic and work now running on other vendors' models, leaving Opus volume where it was. Haiku 5.5 could pull high-volume jobs lower again once it ships [4].
I think the half-price step is nearer a ceiling than a typical saving for teams moving agent work off Opus. Cached context costs the same on both tiers [2], and a model two points behind on GDPval-AA may need extra passes on harder jobs [7]. The case against: if Sonnet finishes routine jobs in fewer tokens than Opus, the saving exceeds half, and same-task logs showing that would prove this view wrong [1].
What to watch
- Haiku 5.5's per-token price, and whether its cache-read rate differs from the $0.20 both larger models charge.
- Independent same-task token counts for Sonnet 5.5 and Opus 5.5; Sonnet needing fewer tokens than Opus would push routing savings past half.
- Any cut to Opus 5.5's $4 and $20 rates, since a narrower gap moves the break-even for sending work down a tier.