Build1 publisher3 min readPublished Updated
Anthropic ships a cache differ, and concedes prompt caching was failing silently
A new beta fingerprints each request and names the first structural divergence from a prior response id. Until now the only signal was cache_read_input_tokens going to zero.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Prompt caching cuts latency and cost significantly, but only when the beginning of your prompt is byte-for-byte identical to a recent request.
- A reordered tool, a timestamp interpolated into your system prompt, or an edit to an earlier message can silently invalidate the cache.
- Without cache diagnostics, the only signal of a cache miss is usage.cache_read_input_tokens dropping to zero, with no indication of what changed.
- Pass the id of your previous response and the API compares the two requests and tells you where they diverged: the model, the system prompt, the tools, or the message history.
- When the beta header is present, the API stores a lightweight fingerprint of each request, keyed by the response id.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Anthropic has documented a beta, `cache-diagnosis-2026-04-07`, in which the API stores a lightweight fingerprint of each request keyed by the response id, then compares a later request against it and attaches a `diagnostics` object naming the first point of divergence [5][6][11]. It matters because the same page concedes that without this, the only signal of a cache miss was `usage.cache_read_input_tokens` dropping to zero, with no indication of what changed [3].
That is the part worth dwelling on. Prompt caching pays only when the beginning of a prompt is byte-for-byte identical to a recent request [1], and the documentation lists the ordinary ways that breaks: a reordered tool, a timestamp interpolated into a system prompt, an edit to an earlier message, each of which can silently invalidate the cache [2]. A silent invalidation shows up as latency and spend, not as an error, so teams have been reverse-engineering their own prompt assembly from a single counter.
The mechanism is a differ, not a cache accountant. You pass the previous response id as `diagnostics.previous_message_id`; the API rebuilds the fingerprint for the new request and compares [6][4]. The reported divergence is one of four categories: the model, the system prompt, the tools, or the message history [4][15]. That is a category, not an offending byte, and the reporting points to a separate response-format section for the possible values rather than enumerating them [12]. Anthropic also states plainly that the comparison is about request structure and is independent of whether the cache actually hit, with a separate section on reading diagnostics alongside `usage` [8]. So this tells you your two requests differ; confirming the cost consequence still means reading the token counters.
The example code has three branches, and the third is the interesting one: `diagnostics` absent means no divergence detected, a `diagnostics` object with `cache_miss_reason` set to null means the comparison is still pending, and otherwise you read `cache_miss_reason.type` [9]. A pending state in a synchronous response path implies the comparison does not always land before the answer does. Treat it as a log line to aggregate, not a guardrail to branch on.
Operationally it is intrusive in a small way: the beta header goes on every turn, the first turn passes `"previous_message_id": null` to opt in with nothing to compare against, and every subsequent turn carries the latest response id forward [10][14]. Streaming users get `diagnostics` on the `message_start` event [13]. On the data side, Anthropic says fingerprints contain only hashes and token-count estimates, never raw prompt content, are retained for a limited time, are scoped to your organization and workspace, and are not used for any other purpose [7].
What to watch. The retention window is described only as "a limited time" and is not quantified in the documentation supplied [16], which matters if you intend to diff a request against something from yesterday's run rather than the previous turn. Watch how often the pending branch fires under real traffic, because a differ that answers late is useful for weekly cost review and useless for a retry decision. And watch whether the four categories get finer: knowing the divergence is in "the message history" narrows a bug in a long agent transcript by very little.