Product1 distinct publisher3 min readPublished
The cache read dropped from 10% of base input to 2.5% while every other rate held, so the saving a team actually books is three quarters of whatever share of last month's bill those reads were.
The Product Desk · Product desk

product
Anthropic bills Pro seats extra for the flagship model already in their picker1 distinct publisher
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 distinct publisher
invest
Open weights take 29% of gateway tokens on a twenty-fifth of the dollars1 distinct publisher
leadership
Every notable AI release today arrived with a grade written by its own vendor1 distinct publisher
Compiled by The Product DeskSomething wrong?How this is made
The cache-read column in the last four weeks of Fable usage, not the figure in the announcement, sets the saving a team actually books.
The rate on that line went from $1.00 per million tokens to $0.25, a cut of 75% [4][13]. Every other cell held [5]. So the fall in your total bill is 0.75 times whatever share of the bill cache reads were. To land Anthropic's typical figure of about 25%, cache reads had to be roughly a third of your Fable spend, since 0.25 divided by 0.75 is 0.33 [14]. To land the up-to-45% figure, roughly 60% [15]. Anthropic describes the agentic case it measured as context-heavy, tool-heavy work in which cache reads make up most of the cost, which is the same statement reached from the other end [12].
Teams often assume that using prompt caching means getting the discount, but the bill works differently. In the shape Anthropic labels long agentic sessions, a large stable prefix is re-sent every turn and grows, so cache reads dominate by construction and the discount lands nearly in full [18]. In mixed production traffic, a cached system prompt plus retrieved context sits alongside a lot of generated output, cache reads are a real but minority share, and the output rate is untouched [20]. For short one-shot calls with no reusable prefix, nothing on the sheet changed, so the saving is zero [16].
Inside Anthropic's own lineup the ordering inverts. Opus 5 charges $0.50 per million for a cache read, twice what Fable 5.1 charges, while costing twice as much on input and output [9][17]. Anthropic's migration guide puts that in front of teams moving up from Opus 5 [10]. The 2x-the-price shorthand stops being a single number and becomes an estimate about session shape, because the longer the session and the larger the prefix, the more of the bill sits on the one line where the flagship is now the cheaper model. Finding the crossover takes your own turn counts and prefix sizes.
There is a second, less visible cost. Anthropic prices cache reads at 10% of base input on every other current Claude model, and OpenAI and Google publish the same multiplier [4][6]. Meta's Contributor tier goes lower at 2.0% in exchange for training permission, and DeepSeek works out near 3.2%, but 2.5% is the lowest ratio a US frontier lab has published [7]. Any forecasting spreadsheet with input divided by ten hard-coded now returns a cached rate four times too high for this one model [8].
The forcing function is one division. Split four weeks of Fable spend into cache reads, fresh input and output, multiply the cache-read share by 0.75, and read the answer as the discount you will book. Above a third, the release is worth the migration on cost grounds alone [14]. At a tenth, the saving is 7.5% and you should be moving for some other reason [19].
Ranked by verification strength, evidence, and original report placement.
Anthropic released Claude Fable 5.1 on September 1, 2026.
Anthropic says Fable 5.1 costs an estimated 25% less than Fable 5 for typical workloads, and up to around 45% less for complex coding and highly agentic tasks.
Anthropic states that Fable 5.1 pricing is otherwise the same as Fable 5's: $10 per million input tokens and $50 per million output tokens.
On every other current Claude model a cache read costs 10% of the base input rate; on Fable 5.1 it costs 2.5%, or $0.25 per million tokens, down from $1.00.
Nothing else on the Fable 5.1 price sheet moved: five of the six cells are identical to Fable 5's, per Anthropic's model documentation.
The common cost-model shortcut of dividing the input rate by ten to get the cached rate now produces a number four times too high on Fable 5.1.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Vendor documents, quoted well, by one reader
Every rate here traces to Anthropic's own model documentation, release post and migration guide, read by a single outlet. Digital Applied quotes them directly, which makes the $0.25 cache read and the untouched $10/$50 pair about as checkable as a solo report gets. The market comparison is a different grade of material: OpenAI and Google at 10%, Meta at 2.0%, DeepSeek near 3.2% arrive with nothing attached. And the piece's own prose calls Opus 5 'half the price' of Fable 5.1 on a line where its figures say double — firm where it quotes, soft where it asserts.
Nobody's bill has been shown
Two days after launch, the only usage signal in this reporting is Anthropic's own retrospective index — four weeks of August 2026 traffic across Claude Enterprise, Claude Code and the API — which measures the workloads used for comparison, not uptake of the new model. No customer deployment, no third-party benchmark, no invoice from a team that actually booked the 25%. What exists is availability across four surfaces and a price sheet.
One cell of six, headlined as the whole sheet
'25% cheaper' is doing more work than the price sheet supports. One of six cells moved, so the saving is three quarters of whatever share cache reads were, and an extraction workload with no reusable prefix saves precisely nothing. Anthropic's own number is qualified and honestly measured — on Anthropic's own surfaces, at Anthropic's chosen effort. Digital Applied is mostly closing that gap rather than widening it, then reopens a smaller one with a market-leadership claim nothing here checks.
Both interests are visible without digging
Anthropic gets to present a single line-item discount as a whole-model price reduction, with the comparison workloads selected, run and indexed in-house. Digital Applied gets a pricing story that routes readers to the per-million-token index it maintains and refreshed the day the model shipped, and a breaking-changes angle that rewards the same expertise. Neither incentive makes a figure wrong; together they explain why the framing lands where it does.
Arithmetic solid, everything comparative unverified
One outlet, two days after the announcement, with the vendor as the ultimate source of every number. The internal arithmetic is checkable and it holds, so the cache-read rate and the share-of-bill logic can be relied on. The rival multipliers, the frontier-lab ranking and the Opus 5 price-order slip are all things a second reading of the same public documentation would settle, and none of them has had one.