Build10 publishersWidely confirmed3 min readPublished
Prompt length sets what Anthropic's Haiku 5.5 price cut is worth to each workload
Anthropic launched Claude Haiku 5.5 at an average price about 75% below Haiku 4.5. The saving varies widely with prompt length, so teams moving classification, support or query traffic need to price their own requests before they switch models.
The Engineer · Build desk

What happened
- Haiku 5.5 ships with an updated tokenizer that Anthropic says consumes slightly more tokens per task than Haiku 4.5's.
- Anthropic halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens and says most agentic tasks get about 20% cheaper.
- Haiku 5.5 is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure under the model ID claude-haiku-5-5.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Long-context users, such as teams running compaction over large histories, lose the most points of discount to any tokenizer overhead and keep the smallest share of the cut.
- decision For pure classification and routing the comparison set is wider than Haiku 4.5: The New Stack lists Qwen3.7 Flash at $0.03 per million input tokens for short inputs, and decision models like Jev lower still.
- decision Agent loops that reread long cached contexts now have a cheaper reason to stay on Sonnet 5.5, so a move to Haiku 5.5 has to beat the post-cut Sonnet bill.
A request-weighted average of Haiku 5.5's two per-token cuts comes to 86% [20]. That uses Anthropic's own split: about 90% of Haiku 4.5 requests fell under the 100,000-token line [12], where the cut is 90%, against 50% above it [3]. Anthropic's published figure is lower, and according to The New Stack it accounts for the request mix and the new tokenizer [6]. A long request carries more tokens than a short one, so the long tier weighs more in a token-weighted bill than its tenth of the request count [20].
Anthropic flagged the tokenizer change itself and called the increase slight [13]. The Decoder notes that the Opus 4.x tokenizer change alone raised token usage by about 30% [26]. Taking 30% as the pessimistic case, a short-tier prompt costs the equivalent of $0.13 per million old-tokenizer tokens, 87% under Haiku 4.5 [21]. A long-tier prompt costs the equivalent of $0.65, a 35% saving [22]. Output lands at the same ratios, because every tier on both models prices output at five times input [22].
The Decoder describes the line as a prompt length [7]. Neither outlet says whether the higher rate covers the whole request or only the tokens past 100,000. If it covers the whole request, a 101,000-token prompt costs about $0.05 in input against about $0.01 for a 99,000-token one [23]. A prompt that sat just under the line on Haiku 4.5's tokenizer can cross it on Haiku 5.5's [13]. Compaction, one of the jobs Anthropic recommends the model for, tends to start from a long context [14].
Haiku 5.5 is also the first Haiku with effort controls, and the setting governs how many tokens it spends on a task; the default is medium [9]. A per-task cost measured at medium is the figure to set against Haiku 4.5 [9]. Raising effort to recover accuracy on hard tickets is paid for in tokens [9].
The quality case rests on Anthropic's own evaluations [10]. Haiku 5.5 scored 72.4% on the offline OSWorld 2.1 subset, against 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna [10]. On Terminal-Bench 4.0 it went from zero to 39.2%, a gain best not quoted as a multiple, while Sonnet 5.5 scored 70.6% [11]. Both suites test computer use and command-line work, and neither scores label accuracy on a support queue [10][11]. I would not move a classification workload on these tables without first scoring a labelled sample of its own traffic at medium effort. Anthropic's guidance keeps Haiku 5.5 on narrowly scoped work and leaves complex agentic coding to Sonnet 5.5 and Opus 5.5 [14].
Anthropic's 20% estimate for Sonnet 5.5 implies that cache reads were about 40% of a typical agentic task's cost before the cut [24]. A loop with a smaller cached share saves less [24]. Haiku 5.5's short-tier input and Sonnet 5.5's cache reads now cost the same $0.10 per million tokens [25]. According to The New Stack, the Sonnet cut rolls out across platforms on launch day, though some existing Azure and Google Cloud customers will wait a few days [16].
What to watch
- Independent token counts for Haiku 5.5 against Haiku 4.5 on identical prompts, to size the tokenizer overhead Anthropic calls slight.
- Artificial Analysis results for Haiku 5.5 on GDPval-AA v2.1, set beside Anthropic's reported 1,620 and the 1,647 it gives Z.ai's GLM-5.3-Flash.
- Whether OpenAI changes pricing on GPT-6 Luna, the budget model Anthropic chose as its comparison.