Skip to content

Topic

Prompt Caching

LLM provider feature reusing computed key-value tensors for repeated prompt prefixes, cutting latency and cost until cache expiry.

Current stories

invest5 publishers

Cached context shrinks the discount from Anthropic's half-price Sonnet 5.5

Anthropic released Sonnet 5.5 at $2 and $10 per million input and output tokens, half the Opus 5.5 rate. How much a buyer saves by moving work down a tier depends on tokens burned per task and on cache reads priced identically on both models.

Perspective Coverage

5 publishers
Builder
Builder 46%
Operator
Operator 38%
Investor
Investor 16%

Reality

Evidence55
Adoption35
Hype gap+20
Incentives70
Confidence60
build1 publisher

Cache expiry decides whether a whole knowledge base in the prompt undercuts retrieval

Cache-augmented generation costs about what retrieval does when the corpus is roughly 10 times the tokens retrieval would send, a dev.to analysis finds. Sparse traffic breaks the rule, because each query then pays the cache-write premium and caching becomes the most expensive option.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence40
build17 publishers

Opus 5.5 matched Opus 5's puzzle answers for up to 69 percent less at its lower default effort

Opus 5.5 matched Opus 5 on two reasoning puzzles in The New Stack's tests at 43 to 69 percent lower cost. Both ran at default effort, medium on the new model and high on the old, so the saving a team sees depends on the effort level it pins.

Perspective Coverage

17 publishers
Builder
Builder 43%
Operator
Operator 33%
Investor
Investor 24%

Reality

Evidence62
Adoption48
Hype gap+22
Incentives58
Confidence58
build1 publisher

One ternary in Jev's gateway limits it to hinting inside Claude Code

Jev's own gateway benchmark shows routing raised Opus 5 input tokens 61% on a Claude Code feature task, where the gateway can only hint at tools. Any saving depends on the task and on how many tools Claude Code sends the router each turn.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+40
Incentives
Insufficient
Confidence40
build14 publishers

Anthropic's Fable 5.1 moves the hard part from prompting to bounding what it may do

The Neuron gave it a browser, a Mac, a broken Blender install and an hour. What changed was how rarely it stopped to ask permission, which makes the next piece of work a harness problem rather than a prompt problem.

Perspective Coverage

14 publishers
Builder
Builder 42%
Operator
Operator 32%
Investor
Investor 26%

Reality

Evidence50
Adoption40
Hype gap+25
Incentives55
Confidence55
product5 publishers

Price cuts minutes apart send agent routing back to the spreadsheet

Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.

Perspective Coverage

5 publishers
Builder
Builder 38%
Operator
Operator 37%
Investor
Investor 25%

Reality

Evidence55
Adoption20
Hype gap+25
Incentives70
Confidence60

Earlier coverage

  1. Each plan-mode toggle under opusplan invalidates Claude Code's prompt cache

    Build · September 23, 2026 · 1 publisher

  2. OpenAI cuts prices on new GPT-6 Sol and Luna models

    Product · September 23, 2026 · 1 publisher

  3. DeepSeek reroutes every V4-Pro API request to V4.1-Flash from 14 September

    Build · September 22, 2026 · 1 publisher

  4. Anthropic's worked example turns a 120,000-token conversation into 2.8 million billed input tokens

    Invest · September 22, 2026 · 1 publisher

  5. A model string one character off bills cached tokens at four times the rate

    Build · September 20, 2026 · 1 publisher

  6. Kimi K3 puts explicit prompt caching on Bedrock behind a 1,024-token minimum prefix

    Build · September 18, 2026 · 1 publisher

  7. A 90% cache-read discount takes 81% off a 10,000-token prompt's input line

    Build · September 17, 2026 · 1 publisher

  8. Ollama divides the whole prompt by the time it spent computing one token of it

    Build · September 15, 2026 · 1 publisher

  9. Bedrock's cache write on the first request pulls the 90 percent discount down to 75

    Build · September 15, 2026 · 1 publisher

  10. A 20-turn agent run bills 656,000 input tokens for 59,000 tokens of reading

    Build · September 15, 2026 · 1 publisher

  11. DeepSeek prices cached agent input at $0.003 a million tokens off-peak

    Product · September 12, 2026 · 1 publisher

  12. A timestamp in the system prompt turns prompt caching into a 25% surcharge

    Build · September 12, 2026 · 1 publisher

  13. A top-level cache_control field moves the Claude cache breakpoint forward as the conversation grows

    Build · September 11, 2026 · 1 publisher

  14. One API key per platform puts every customer's system prompt in the same cache namespace

    Build · September 5, 2026 · 1 publisher

  15. Customer name placed ~1,000 tokens in coincided with a common 1,024-token cache floor in call logs

    Build · September 4, 2026 · 1 publisher

  16. Anthropic's 25% cheaper Fable 5.1 discounts one of six lines on the price sheet

    Product · September 3, 2026 · 1 publisher

  17. Compaction that cut tool output 38.4% pushed the bill up 6.8%

    Build · September 2, 2026 · 1 publisher

  18. DeepSeek V4 moves the coding-model decision into the finance column

    Build · September 1, 2026 · 1 publisher

  19. Per-PTU throughput spans 25x across three models in the same GPT-5.6 family

    Build · August 30, 2026 · 1 publisher

  20. A semantic cache hit saves five times what a prompt cache hit saves

    Build · August 29, 2026 · 1 publisher

  21. A cached prompt prefix repays its write premium on the second request

    Build · August 28, 2026 · 1 publisher

  22. Anthropic's GA Files API re-bills the whole document on every request

    Build · August 27, 2026 · 1 publisher

  23. Claude's limits are token meters on two clocks, and your open session is what drains them

    Build · August 25, 2026 · 1 publisher

  24. Coding agents cost $4,125 a month because 73% of it is context you already sent

    Build · August 23, 2026 · 1 publisher

  25. Tier the models; the validation boundary is the thing you are actually buying

    Build · August 22, 2026 · 1 publisher

  26. Claude's prompt cache dies quietly in agent loops: the 20-block lookback nobody configures

    Build · August 21, 2026 · 1 publisher

  27. Anthropic ships a cache differ, and concedes prompt caching was failing silently

    Build · August 17, 2026 · 1 publisher

  28. DeepSeek's 12x cached-token rise ends the cheap-endpoint era for Chinese inference

    Invest · August 17, 2026 · 1 publisher

  29. Prompt caching cuts agent API costs 41-80%, but only if tool results stay out of the cache

    Build · August 16, 2026 · 1 publisher