Product2 publishers3 min readPublished
Box measured Claude Opus 5.5 using a third of the tokens Opus 5 needed
Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.
The Product Desk · Product desk

What happened
- Anthropic released Claude Opus 5.5 less than two months after Opus 5, and says Sonnet 5.5 and Haiku 5.5 will launch over the coming weeks.
- Tokens on Opus 5.5 are priced 20 percent below Opus 5, and Anthropic says the model needs fewer tokens for higher quality work.
- When safeguards fire, requests drop from Opus 5.5 to Opus 4.8, and most cybersecurity tasks are rerouted to Opus 4.8 while biology work goes to Opus 5.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- cost A quarterly agent forecast built on Opus 5 token volume is off by close to three times if the Box ratio holds on your own content, and the team that committed the number is the one explaining the variance.
- decision The saving big enough to justify re-baselining rests on one customer's evaluation of its own corpus, so a buyer quoting it in a business case is quoting Box's workload until they reproduce it.
- constraint Pipelines that touch security triage cannot be modelled at a single token rate, because part of the traffic lands on Opus 4.8 and part on Opus 5.
- capability A $20-a-month subscriber gets roughly half again as much run capacity inside a five-hour window, which changes how much one person can finish before the limit resets.
Box's token ratio speaks to both the billing and the reading, and it is the hardest figure in this launch to check. Verbose output gets billed twice, once as output tokens and once as the minutes a reviewer spends scrolling to the part that answers the question, and only the tokens show up in a cost model.
Two numbers in the coverage both read 40 percent and count different things. ZDNET's framing is that Opus 5.5 delivers Fable 5.1 performance for most work and costs about 40 percent less to run [2]. Yashodha Bhavnani, VP of AI Products at Box, said that in the company's evaluations Opus 5.5 "used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy" [5].
The published price is the part a buyer can hold Anthropic to: tokens on Opus 5.5 cost 20 percent less than on Opus 5 [3]. Stack that on Box's ratio and the same task bills at 0.8 times a third of the old token count, about 27 percent of the Opus 5 cost, a reduction near 73 percent [19]. The price cut accounts for 20 points of that, and the token ratio for the other 53 [22].
The write-up does not split Box's token count into input and output. If most of the saving sits on the output side, a job that reads a large corpus and writes a short answer sees less of it. Bhavnani said the improvement should "matter a lot for teams running agents across their content in areas like financial services and the public sector" [7].
The other named customer is talking about how the output reads. John Ruelas, a staff software engineer at Ramp, said "Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it" [12]. He said the model "writes like a good colleague and follows our writing rules" [13]. Anthropic separately says Opus 5.5 generates output more than 30% faster than Opus 5 [8].
The efficiency stops at the safeguard boundary. When safeguards fire, requests fall back from Opus 5.5 to Opus 4.8, and ZDNET reports that in practice most cybersecurity tasks get rerouted to Opus 4.8, with biology and frontier LLM development requests sent to Opus 5 [15]. So a triage pipeline cannot be priced at one blended token rate. Organisations approved through Anthropic's Cyber Verification Program can start using Opus 5.5 in a few weeks, and ZDNET's writer disclosed that he has been approved into that program himself [16][18].
For the $20 a month tier the two changes compound. Anthropic raised the five-hour usage limits by 20 percent [9] and says both five-hour and weekly limits go further because Opus 5.5 costs less than Opus 5 [10]. ZDNET pairs a 20 percent larger bucket with a 25 percent slower burn and puts effective run capacity about 50 percent higher [11]; 1.2 times 1.25 is 1.5 [20]. That only holds if the limit is metered in something that tracks the token price.
Running the three jobs that account for most of your token spend on both models, counting input and output tokens separately, costs an afternoon. Two numbers come out of that comparison and only one is safe to put in a budget: 20 percent off per token is published pricing, and a third of the tokens is Box's result on Box's content.
What to watch
- Whether Sonnet 5.5 and Haiku 5.5 ship with the same token-efficiency claim attached, since cheaper tiers carry most agent volume.
- Whether any buyer publishes a token-count comparison with the evaluation method and the input/output split attached.
- How much of a security pipeline actually lands on Opus 4.8 once Cyber Verification Program customers get access.