Build1 publisher2 min readPublished
Cache expiry decides whether a whole knowledge base in the prompt undercuts retrieval
Cache-augmented generation costs about what retrieval does when the corpus is roughly 10 times the tokens retrieval would send, a dev.to analysis finds. Sparse traffic breaks the rule, because each query then pays the cache-write premium and caching becomes the most expensive option.
The Engineer · Build desk

What happened
- Cached input tokens cost about 0.1 times normal input at Anthropic and, on current models, OpenAI, while writing the cache costs about 1.25 times, once.
- The implementation is a prompt layout: the whole corpus first in a fixed byte-for-byte order, a cache breakpoint at its end, and the question after it.
- In the post's worked example, retrieval still wins on tokens for a 200,000-token corpus, while caching the whole thing is cheaper at 20,000 tokens.
- The post keeps retrieval for large corpora, fast-changing content, per-user permissions and provenance.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A service whose queries arrive further apart than the cache lifetime pays 1.25 times the corpus on every call, a quarter more than sending the corpus with no caching at all.
- decision With flagship context windows at a million tokens, the build decision moves from whether the corpus fits to how closely queries are spaced.
- constraint Every edit to the cached corpus invalidates the cache for all later requests, so a frequently edited knowledge base pays the write premium again after each change.
The 10x figure is a warm-cache rule. It prices one cached read of the whole corpus against one uncached retrieval payload [2]. The write premium sits outside that comparison. It is billed each time the provider has to populate the cache [1].
Take a 20,000-token handbook at the post's assumed $5 per million input tokens, with cached reads at about $0.50 per million [3]. A warm query costs about $0.01 in corpus tokens [1]. A query that has to write the cache costs about $0.125 [2]. The miss costs 12.5 times the hit [3].
According to the post, cache entries expire in minutes, and the author names the time-to-live as the real constraint, ahead of the context window [5]. Queries spaced further apart than that each pay the write. "The question that actually decides it is how often you ask," the author wrote [15].
Between warm and cold, count hits per cache lifetime. Call the corpus C tokens, count corpus and retrieval tokens only, and take a corpus five times the retrieval payload. A window of n queries costs 1.25C for the first call and 0.1C for each one after it, against 0.2C per call for retrieval. Caching pulls ahead at about 12 queries per window [5]. At exactly 10x the write never pays back. Caching costs 1.15C more than retrieval in every window, whatever the volume [6]. The author's case for the tie rests on operations. By the post's account, the token comparison leaves out the vector database, the embedding model called on every write and query, chunk tuning, the reranker and the eval harness [13].
The design composes cleanly with response caching. A response cache in front answers repeated identical questions, and the prompt cache behind it serves new questions over the same documents [16]. The author's Java sample sets the cache lifetime in code, with `.ttl(CacheControlEphemeral.Ttl.TTL_1H)` on the system block that holds the corpus [11]. The post does not give a price for the one-hour setting.
Below a provider's minimum prefix, caching is skipped without an error and the corpus bills at full price, forever if nobody looks [9]. The post lists minimums of 512 to 4,096 tokens at Anthropic depending on model, 1,024 on OpenAI's GPT-5.6 and later, and 2,048 to 4,096 at Gemini [9]. The sample logs cache-write, cache-read and uncached token counts on each response, under the comment "The only proof that caching is working." [12]
What to watch
- Any change to the 0.1x cached-read or 1.25x write multipliers at Anthropic or OpenAI, since the 10x break-even comes straight from the read price.
- Published pricing for the one-hour TTL the author's sample requests; that price sets whether low-traffic services can keep a cache warm between queries.