Skip to content

Build1 publisher3 min readPublished Updated

Anthropic ships a cache differ, and concedes prompt caching was failing silently

A new beta fingerprints each request and names the first structural divergence from a prior response id. Until now the only signal was cache_read_input_tokens going to zero.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Anthropic ships a cache differ, and concedes prompt caching was failing silently
Generated illustration

What happened

  • Prompt caching cuts latency and cost significantly, but only when the beginning of your prompt is byte-for-byte identical to a recent request.
  • A reordered tool, a timestamp interpolated into your system prompt, or an edit to an earlier message can silently invalidate the cache.
  • Without cache diagnostics, the only signal of a cache miss is usage.cache_read_input_tokens dropping to zero, with no indication of what changed.
  • Pass the id of your previous response and the API compares the two requests and tells you where they diverged: the model, the system prompt, the tools, or the message history.
  • When the beta header is present, the API stores a lightweight fingerprint of each request, keyed by the response id.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Anthropic has documented a beta, `cache-diagnosis-2026-04-07`, in which the API stores a lightweight fingerprint of each request keyed by the response id, then compares a later request against it and attaches a `diagnostics` object naming the first point of divergence [5][6][11]. It matters because the same page concedes that without this, the only signal of a cache miss was `usage.cache_read_input_tokens` dropping to zero, with no indication of what changed [3].

That is the part worth dwelling on. Prompt caching pays only when the beginning of a prompt is byte-for-byte identical to a recent request [1], and the documentation lists the ordinary ways that breaks: a reordered tool, a timestamp interpolated into a system prompt, an edit to an earlier message, each of which can silently invalidate the cache [2]. A silent invalidation shows up as latency and spend, not as an error, so teams have been reverse-engineering their own prompt assembly from a single counter.

The mechanism is a differ, not a cache accountant. You pass the previous response id as `diagnostics.previous_message_id`; the API rebuilds the fingerprint for the new request and compares [6][4]. The reported divergence is one of four categories: the model, the system prompt, the tools, or the message history [4][15]. That is a category, not an offending byte, and the reporting points to a separate response-format section for the possible values rather than enumerating them [12]. Anthropic also states plainly that the comparison is about request structure and is independent of whether the cache actually hit, with a separate section on reading diagnostics alongside `usage` [8]. So this tells you your two requests differ; confirming the cost consequence still means reading the token counters.

The example code has three branches, and the third is the interesting one: `diagnostics` absent means no divergence detected, a `diagnostics` object with `cache_miss_reason` set to null means the comparison is still pending, and otherwise you read `cache_miss_reason.type` [9]. A pending state in a synchronous response path implies the comparison does not always land before the answer does. Treat it as a log line to aggregate, not a guardrail to branch on.

Operationally it is intrusive in a small way: the beta header goes on every turn, the first turn passes `"previous_message_id": null` to opt in with nothing to compare against, and every subsequent turn carries the latest response id forward [10][14]. Streaming users get `diagnostics` on the `message_start` event [13]. On the data side, Anthropic says fingerprints contain only hashes and token-count estimates, never raw prompt content, are retained for a limited time, are scoped to your organization and workspace, and are not used for any other purpose [7].

What to watch. The retention window is described only as "a limited time" and is not quantified in the documentation supplied [16], which matters if you intend to diff a request against something from yesterday's run rather than the previous turn. Watch how often the pending branch fires under real traffic, because a differ that answers late is useful for weekly cost review and useless for a retry decision. And watch whether the four categories get finer: knowing the divergence is in "the message history" narrows a bug in a long agent transcript by very little.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories