Build1 publisher3 min readPublished
One microsecond timestamp kept a Claude app's prompt cache at zero hits for 3,412 calls
Preterview's Claude calls read nothing from the prompt cache across 3,412 requests because a microsecond timestamp opened the system prompt. Each call paid the 1.25x write price for a prefix nobody read back, so the bill came out higher than with caching off.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Claude's cache matches only an exact byte-for-byte prefix, hashed as tools, then system, then messages, so a change to one early character makes everything after it miss.
- The only evidence sat in the usage fields, where input_tokens looked small because it counts just the uncached remainder, not the total.
- Moving the elapsed time into the latest user message lifted the hit rate to 78%, where it stalled.
- The tool definitions came from a Python set whose order differs per worker process, so four uvicorn workers sent four different prefixes.
- Sorting the tool list raised the hit rate to 93%, and the developer puts the cached prefix at about 88% cheaper per session than under the bug.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Turning caching on without checking hits can cost 25% more per resent prefix token than leaving it off, and at Preterview it ran for nine days before anyone looked.
- constraint Clocks and any other per-call value have to go after the last cache_control marker; Preterview's elapsed time now travels in the newest user turn.
- exposure Python services running several worker processes can send a different prefix from each worker whenever a set feeds the request, and a single-worker test will not show it.
Preterview's system prompt opened with the line `Current time: {datetime.now().isoformat()}`, placed ahead of the persona, rubric and resume [3]. The isoformat() string includes microseconds, so no two requests shared a first line [3]. According to the developer, a request gets a cache read only when its exact prefix was written in the last five minutes. Anything else is a write that starts a new five-minute entry [5]. The clock was there so the interviewer could pace itself [19]. Pacing an interview does not need microseconds. "Reasonable feature. Terrible placement," the developer wrote [10].
Writes are billed at 1.25x the normal input price, reads at 0.1x and uncached input at 1x [6]. An average session runs 22 calls, and each one resends about 6,500 tokens of prefix [9]. Under the bug every call paid the write rate [2]. That put a session's prefix at 27.5 units, against 22 with caching switched off [1]. One write and 21 reads cost 3.35 units, about 12% of the buggy figure [1]. The developer estimates the cached prefix now costs about 88% less per session after both fixes, which matches that calculation [15]. The discount is big enough that one read per write already comes out ahead: two calls cost 1.35 units cached against 2.0 uncached [5]. Money is lost only on entries nobody reads back [7]. With a first line that changed on every call, every entry went unread, about 22 million prefix tokens across the 3,412 calls [3].
Nothing looked broken. Responses and latency were normal [11]. One logged call shows 212 input tokens, 6,488 cache-creation tokens and zero cache reads [11]. On the monthly invoice, the post says, the fault appeared only as "input tokens, slightly more than expected" [20]. "The API never warns you. The only signal is the usage fields. Log them," the developer wrote [16]. Preterview now emits a metric on every call, read tokens divided by read plus write plus fresh tokens, tagged by session and interviewer style [17].
The second bug showed up only with more than one worker process [14]. Tools are hashed before the system prompt and messages [4]. So when a worker's tool order differed, the match also failed on the persona, rubric and resume that follow the tools. A session spread over four workers can pay up to four writes instead of one [4]. Four writes in 22 calls leave 18 reads, about 82%, close to the 78% plateau [4].
The repaired request is careful about order. The system prompt is now two cached blocks: persona and rubric first, the candidate's resume second [18]. The elapsed time goes at the end of the newest user message as `[elapsed: {elapsed_min} of 30 min]` [18]. The first block depends only on which of the three interviewer styles is running, so it changes less often than the resume [8][18]. I think that ordering, stable content first and per-call values after the last marker, is the right default for any chat loop that resends its context.
The 88% figure only carries over to traffic shaped like Preterview's [15]. Twenty-two calls in 30 minutes is one call about every 1.4 minutes, well inside a five-minute lifetime that each hit refreshes [2]. A workload that comes back to the same prefix less than once every five minutes will write on most calls and see little of that saving [5].
What to watch
- Hit-rate data from workloads with longer gaps between calls than Preterview's roughly 1.4 minutes, where the five-minute cache lifetime can lapse between turns.
- Any API change that flags cache entries written but never read; for now the usage fields are the only signal.
- Whether new tool definitions or interviewer styles keep Preterview's tools block byte-stable, since it hashes ahead of everything else.