Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Anthropic's October 7 price cut halves one line of the Sonnet 5.5 bill

Anthropic cut the Claude Sonnet 5.5 cache-read rate from $0.20 to $0.10 per million tokens on October 7. A team saves $0.10 for each million cache-read tokens it logs, so its total bill drops by half the share those reads already held.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Anthropic's October 7 price cut halves one line of the Sonnet 5.5 bill
Generated illustration

What happened

  • The API's response usage object reports cache-read tokens separately from cache-creation and ordinary input tokens, so each call's read volume can be logged.
  • In a worked example from a dev.to post, 25 million cache-read tokens a month cost $5 at the old rate and $2.50 at the new one.
  • The post advises grouping traffic by model ID and operation, because mixing models or request types can make a workload shift look like a price effect.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A team's total Sonnet 5.5 bill shrinks by half its pre-cut cache-read share, so routes dominated by output or uncached input get a small fraction of the headline 50 percent.
  • decision Adding caching to a prompt that varies per call still incurs writes at the unchanged price, and without enough later reads that route can end up costing more.
  • constraint A defensible before-and-after figure depends on counters split by model and route; a pooled invoice total cannot separate the price cut from a change in traffic mix.

The saving scales with one counter. A dev.to post on the change puts the direct difference at $0.10 per million cache-read tokens, assuming the same volume and billing terms [3]. Its formula for an existing workload takes recorded cache-read tokens, divides by one million, and multiplies by the gap between the old and new rates [4]. Each dollar saved takes 10 million cache-read tokens [17]. A route has to serve a billion cache-read tokens a month before the cut is worth $100 a month on it [18].

Measured against the whole invoice, the saving gets smaller. Output generation, uncached input, cache creation, retries and other models are separate lines [14]. The post quotes Anthropic's release notes as saying cache writes and all other prices are unchanged [2]. A total bill therefore falls by half of whatever fraction of it was cache-read spend before the October 7 cut [16]. The full 50 percent would need a bill made only of reads. That is an odd workload, given that the first request to a prefix has to write the cache before anything can read it [6]. The post does not quote the write or uncached-input rates.

Writes are where the cut can mislead. A prompt that changes on every call can create cache writes without earning enough later reads to offset them, and writes are still billed at the old rate [6]. The post's advice is to put stable content ahead of request-specific data and confirm that the exact route reuses that prefix [10]. The routes that gain already hit a large repeated prefix, such as a stable system instruction, shared reference material or tool definitions [9]. A larger read saving can also sit beside higher write or uncached-input spend on the same route [12].

Measuring it is cheap because the API already splits the counters. The response usage object reports cache_read_input_tokens separately from cache_creation_input_tokens and ordinary input_tokens [7]. That is good API design. The number that decides the saving comes back with every response, and no prompt text has to be stored to use it. The post recommends persisting those counters with model ID, route, timestamp, request outcome and an internal request class, while keeping customer prompts and secrets out of the log [11].

Before-and-after comparisons usually go wrong at the grouping step. Pool Sonnet 5.5 with other models, or pool unrelated request types, and a shift in workload looks like a price effect [8]. The post's checklist for each comparison window covers completed calls, the four token counters, errors and retries. It also reports per-call averages beside totals when traffic volume differs [13].

I think the post has the order right. Compute the read delta from logged counters and treat it as an estimate of one component at published rates. Keep the raw counts, so the figure can be recomputed if a rate or a negotiated billing term changes [15].

What to watch

  • Any Anthropic change to Sonnet 5.5 cache-write or uncached-input pricing; the October 7 note left both unchanged.
  • First post-cut invoices checked against logged counters: a billed read line that differs from $0.10 per million tokens points to negotiated terms.
  • Whether Anthropic extends the read cut to other Claude models, whose prices the release notes left unchanged.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption
Insufficient
Hype gap0
Incentives
Insufficient
Confidence60
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Anthropic cut the Claude Sonnet 5.5 prompt-cache read rate on October 7, 2026, from $0.20 to $0.10 per million tokens.

    ReportedSupportedSource: dev.to post citing Anthropic release notesView cited source
  2. [2]

    Anthropic's release notes say Sonnet 5.5 cache writes and all other prices are unchanged.

    ReportedSupportedSource: dev.to post citing Anthropic release notesView cited source
  3. [3]

    The direct difference is $0.10 for each million cache-read tokens, assuming the same token volume and billing terms.

    ReportedSupportedSource: dev.to postView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 8, 2026

    Claude Sonnet 5.5 cache reads cost half: measure your own savings

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories