Build1 publisher3 min readPublished
Prompt caching cuts agent API costs 41-80%, but only if tool results stay out of the cache
An arXiv evaluation across OpenAI, Anthropic and Google reports that naive full-context caching can raise latency, while excluding dynamic tool results gives more consistent gains.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A paper titled "Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks" is posted on arXiv with identifier 2601.06007, listed under Computer Science > Computation and Language.
- The authors state that to their knowledge no prior work quantifies the cost benefits of prompt caching or compares caching strategies for multi-turn agentic tasks.
- The evaluation covers three LLM providers (OpenAI, Anthropic and Google) using four flagship models.
- Three caching strategies were compared: full context caching, system prompt only caching, and caching that excludes dynamic tool results.
- Evaluation used DeepResearch Bench, a multi-turn agentic benchmark where agents autonomously execute real-world web search tool calls to answer complex research questions.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A preprint posted to arXiv under Computation and Language, identifier 2601.06007, measures what prompt caching actually does to cost and latency in long-horizon agent runs across OpenAI, Anthropic and Google [1]. The headline result is that caching cut API costs by 41-80% and improved time to first token by 13-31% across providers, but the authors report that the crude version of the technique, caching the whole context, can paradoxically increase latency [7][8].
The setup matters for whether you believe the numbers. The authors evaluated four flagship models across the three providers [3], on DeepResearch Bench, where agents autonomously execute real web search tool calls to answer research questions [5], across more than 500 agent sessions with 10,000-token system prompts, measuring API cost and TTFT [6]. Three strategies were compared: full context caching, system prompt only caching, and caching that excludes dynamic tool results [4]. Prompt caching itself works by reusing previously computed key-value tensors from attention layers so that repeated prompt prefixes are not recomputed [10], which is why anything dynamic sitting early in the prompt is expensive: it invalidates everything behind it.
That is the operational finding. The paper reports that deliberate cache block control gives more consistent benefits than full-context caching, specifically placing dynamic content at the end of the system prompt, avoiding dynamic traditional function calling, and excluding dynamic tool results from the cached region [8]. In other words, the win comes from prompt layout discipline, not from turning a feature on. Agent frameworks that append tool output into the shared prefix are the ones most likely to see caching underdeliver, since long-horizon conversations accumulate dozens of API calls and tens of thousands of tokens [15].
An ablation over prompt sizes from 500 to 50,000 tokens and tool call counts from 3 to 50 found linear cost and TTFT benefits once the provider's minimum cacheable token count is exceeded, and also surfaced provider-specific discrepancies between strategy variants [9]. Two consequences for anyone budgeting: below the minimum you get nothing, and a strategy tuned on one provider is not guaranteed to transfer to another.
Caveats. The reported ranges are wide, with the top of the cost saving band nearly twice the bottom [14], and the abstract does not break the figures out by provider or model [16], so a single-provider planning number cannot be taken from it. This is one preprint, and the authors describe it as the first work to quantify these savings or compare caching strategies for multi-turn agentic tasks [2]; prior literature they cite addresses inference-level KV cache memory management and compression rather than provider API caching features [11], and concurrent work audited provider prompt caching for timing side-channel exposure [12]. The keyword block lists PricewaterhouseCoopers, U.S. [13].
Worth watching: whether the per-provider breakdowns and the minimum cacheable token thresholds hold as providers change pricing, and whether agent frameworks start treating tool-result placement as a first-class configuration rather than an implementation detail [8][9].