Build1 publisherNot yet confirmed elsewhere3 min readPublished
Anthropic's October 7 price cut halves one line of the Sonnet 5.5 bill
Anthropic cut the Claude Sonnet 5.5 cache-read rate from $0.20 to $0.10 per million tokens on October 7. A team saves $0.10 for each million cache-read tokens it logs, so its total bill drops by half the share those reads already held.
The Engineer · Build desk

What happened
- The API's response usage object reports cache-read tokens separately from cache-creation and ordinary input tokens, so each call's read volume can be logged.
- In a worked example from a dev.to post, 25 million cache-read tokens a month cost $5 at the old rate and $2.50 at the new one.
- The post advises grouping traffic by model ID and operation, because mixing models or request types can make a workload shift look like a price effect.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A team's total Sonnet 5.5 bill shrinks by half its pre-cut cache-read share, so routes dominated by output or uncached input get a small fraction of the headline 50 percent.
- decision Adding caching to a prompt that varies per call still incurs writes at the unchanged price, and without enough later reads that route can end up costing more.
- constraint A defensible before-and-after figure depends on counters split by model and route; a pooled invoice total cannot separate the price cut from a change in traffic mix.
The saving scales with one counter. A dev.to post on the change puts the direct difference at $0.10 per million cache-read tokens, assuming the same volume and billing terms [3]. Its formula for an existing workload takes recorded cache-read tokens, divides by one million, and multiplies by the gap between the old and new rates [4]. Each dollar saved takes 10 million cache-read tokens [17]. A route has to serve a billion cache-read tokens a month before the cut is worth $100 a month on it [18].
Measured against the whole invoice, the saving gets smaller. Output generation, uncached input, cache creation, retries and other models are separate lines [14]. The post quotes Anthropic's release notes as saying cache writes and all other prices are unchanged [2]. A total bill therefore falls by half of whatever fraction of it was cache-read spend before the October 7 cut [16]. The full 50 percent would need a bill made only of reads. That is an odd workload, given that the first request to a prefix has to write the cache before anything can read it [6]. The post does not quote the write or uncached-input rates.
Writes are where the cut can mislead. A prompt that changes on every call can create cache writes without earning enough later reads to offset them, and writes are still billed at the old rate [6]. The post's advice is to put stable content ahead of request-specific data and confirm that the exact route reuses that prefix [10]. The routes that gain already hit a large repeated prefix, such as a stable system instruction, shared reference material or tool definitions [9]. A larger read saving can also sit beside higher write or uncached-input spend on the same route [12].
Measuring it is cheap because the API already splits the counters. The response usage object reports cache_read_input_tokens separately from cache_creation_input_tokens and ordinary input_tokens [7]. That is good API design. The number that decides the saving comes back with every response, and no prompt text has to be stored to use it. The post recommends persisting those counters with model ID, route, timestamp, request outcome and an internal request class, while keeping customer prompts and secrets out of the log [11].
Before-and-after comparisons usually go wrong at the grouping step. Pool Sonnet 5.5 with other models, or pool unrelated request types, and a shift in workload looks like a price effect [8]. The post's checklist for each comparison window covers completed calls, the four token counters, errors and retries. It also reports per-call averages beside totals when traffic volume differs [13].
I think the post has the order right. Compute the read delta from logged counters and treat it as an estimate of one component at published rates. Keep the raw counts, so the figure can be recomputed if a rate or a negotiated billing term changes [15].
What to watch
- Any Anthropic change to Sonnet 5.5 cache-write or uncached-input pricing; the October 7 note left both unchanged.
- First post-cut invoices checked against logged counters: a billed read line that differs from $0.10 per million tokens points to negotiated terms.
- Whether Anthropic extends the read cut to other Claude models, whose prices the release notes left unchanged.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence60
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Anthropic cut the Claude Sonnet 5.5 prompt-cache read rate on October 7, 2026, from $0.20 to $0.10 per million tokens.
- [2]
Anthropic's release notes say Sonnet 5.5 cache writes and all other prices are unchanged.
- [3]
The direct difference is $0.10 for each million cache-read tokens, assuming the same token volume and billing terms.
- [4]
For an existing workload: read_cost_delta = cache_read_tokens / 1,000,000 x ($0.20 - $0.10).
- [5]
In a hypothetical month with 25 million Sonnet 5.5 cache-read tokens, that read volume cost $5 at the former rate and $2.50 at the new rate.
- [6]
A first request can create cached content, and cache writes are unchanged by the price change; a prompt that changes on every call may create writes without earning enough later reads to offset them.
- [7]
The response usage object distinguishes cache reads (cache_read_input_tokens) from cache creation (cache_creation_input_tokens) and ordinary input (input_tokens), alongside output_tokens.
- [8]
Group traffic by model ID and operation; combining Sonnet 5.5 with other models, or unrelated request types, can make a workload shift look like a price effect.
- [9]
The change matters most to routes that already get cache hits on a substantial repeated prefix, such as a stable system instruction, shared reference material, or tool definitions reused across requests.
- [10]
Put stable content before request-specific data, and measure whether the exact route is reusing that prefix.
- [11]
Persist usage counters with model ID, route, timestamp, request outcome and an internal request class; avoid storing customer prompts or secrets, since aggregated counters are enough for the price comparison.
- [12]
A larger read saving can coexist with higher write or uncached-input spend.
- [13]
For each comparison window, record completed calls, cache-read tokens, cache-creation tokens, ordinary input tokens, output tokens, errors and retries; if traffic volume differs, report both totals and per-call averages.
- [14]
The cut is not a claim that total Anthropic spend falls by 50 percent: output generation, uncached input, cache creation, request volume, retries and other models remain separate parts of the bill.
- [15]
The calculation is an estimate of the cache-read component at published rates, not a substitute for the billed total; retain raw counters so the comparison can be recomputed if the rate or a negotiated billing term changes.
- [16]
Because only the read rate moved, the fraction by which a total Sonnet 5.5 bill falls equals half of the share of that bill that was cache-read spend before the cut.
- [17]
Each $1 of saving requires 10 million cache-read tokens.
- [18]
A $100 monthly saving requires 1 billion cache-read tokens a month.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toClaude Sonnet 5.5 cache reads cost half: measure your own savings
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.