Skip to content

Build10 publishersWidely confirmed3 min readPublished

Prompt length sets what Anthropic's Haiku 5.5 price cut is worth to each workload

Anthropic launched Claude Haiku 5.5 at an average price about 75% below Haiku 4.5. The saving varies widely with prompt length, so teams moving classification, support or query traffic need to price their own requests before they switch models.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Prompt length sets what Anthropic's Haiku 5.5 price cut is worth to each workload
Generated illustration

What happened

  • Haiku 5.5 ships with an updated tokenizer that Anthropic says consumes slightly more tokens per task than Haiku 4.5's.
  • Anthropic halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens and says most agentic tasks get about 20% cheaper.
  • Haiku 5.5 is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure under the model ID claude-haiku-5-5.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Long-context users, such as teams running compaction over large histories, lose the most points of discount to any tokenizer overhead and keep the smallest share of the cut.
  • decision For pure classification and routing the comparison set is wider than Haiku 4.5: The New Stack lists Qwen3.7 Flash at $0.03 per million input tokens for short inputs, and decision models like Jev lower still.
  • decision Agent loops that reread long cached contexts now have a cheaper reason to stay on Sonnet 5.5, so a move to Haiku 5.5 has to beat the post-cut Sonnet bill.

A request-weighted average of Haiku 5.5's two per-token cuts comes to 86% [20]. That uses Anthropic's own split: about 90% of Haiku 4.5 requests fell under the 100,000-token line [12], where the cut is 90%, against 50% above it [3]. Anthropic's published figure is lower, and according to The New Stack it accounts for the request mix and the new tokenizer [6]. A long request carries more tokens than a short one, so the long tier weighs more in a token-weighted bill than its tenth of the request count [20].

Anthropic flagged the tokenizer change itself and called the increase slight [13]. The Decoder notes that the Opus 4.x tokenizer change alone raised token usage by about 30% [26]. Taking 30% as the pessimistic case, a short-tier prompt costs the equivalent of $0.13 per million old-tokenizer tokens, 87% under Haiku 4.5 [21]. A long-tier prompt costs the equivalent of $0.65, a 35% saving [22]. Output lands at the same ratios, because every tier on both models prices output at five times input [22].

The Decoder describes the line as a prompt length [7]. Neither outlet says whether the higher rate covers the whole request or only the tokens past 100,000. If it covers the whole request, a 101,000-token prompt costs about $0.05 in input against about $0.01 for a 99,000-token one [23]. A prompt that sat just under the line on Haiku 4.5's tokenizer can cross it on Haiku 5.5's [13]. Compaction, one of the jobs Anthropic recommends the model for, tends to start from a long context [14].

Haiku 5.5 is also the first Haiku with effort controls, and the setting governs how many tokens it spends on a task; the default is medium [9]. A per-task cost measured at medium is the figure to set against Haiku 4.5 [9]. Raising effort to recover accuracy on hard tickets is paid for in tokens [9].

The quality case rests on Anthropic's own evaluations [10]. Haiku 5.5 scored 72.4% on the offline OSWorld 2.1 subset, against 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna [10]. On Terminal-Bench 4.0 it went from zero to 39.2%, a gain best not quoted as a multiple, while Sonnet 5.5 scored 70.6% [11]. Both suites test computer use and command-line work, and neither scores label accuracy on a support queue [10][11]. I would not move a classification workload on these tables without first scoring a labelled sample of its own traffic at medium effort. Anthropic's guidance keeps Haiku 5.5 on narrowly scoped work and leaves complex agentic coding to Sonnet 5.5 and Opus 5.5 [14].

Anthropic's 20% estimate for Sonnet 5.5 implies that cache reads were about 40% of a typical agentic task's cost before the cut [24]. A loop with a smaller cached share saves less [24]. Haiku 5.5's short-tier input and Sonnet 5.5's cache reads now cost the same $0.10 per million tokens [25]. According to The New Stack, the Sonnet cut rolls out across platforms on launch day, though some existing Azure and Google Cloud customers will wait a few days [16].

What to watch

  • Independent token counts for Haiku 5.5 against Haiku 4.5 on identical prompts, to size the tokenizer overhead Anthropic calls slight.
  • Artificial Analysis results for Haiku 5.5 on GDPval-AA v2.1, set beside Anthropic's reported 1,620 and the 1,647 it gives Z.ai's GLM-5.3-Flash.
  • Whether OpenAI changes pricing on GPT-6 Luna, the budget model Anthropic chose as its comparison.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories