Build1 publisher3 min readPublished
Your 90% Cache Hit Ratio Is a Lagging Indicator. Alert on Cold Misses Per Key
A small FastAPI repro turns 20 simultaneous requests for one tenant into 20 database queries while the dashboard still reads healthy. The metric that catches it is in-flight loads per key.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- While profiling a small FastAPI service, the author found single requests were fast, the cache hit ratio sat above 90 percent, and the database showed only a modest load during ordinary traffic.
- The problem appeared only when a release or a retry storm sent twenty requests for the same tenant at once; database CPU climbed while the cache hit ratio stayed high enough to look innocent.
- The service cached tenant project lists for thirty seconds.
- A high hit ratio only tells you what happened after a key was already warm; it does not tell you how many cold-start requests turned into database work at the same moment.
- In the minimal reproduction, a ThreadPoolExecutor with 20 workers issues 20 requests to /projects/tenant-a and the output shows db_calls: 20.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer profiling a small FastAPI service found the cache hit ratio sitting above 90 percent, single requests fast, and only modest database load during ordinary traffic [1]. The database CPU still climbed whenever a release or a retry storm sent twenty requests for the same tenant at once, and the hit ratio stayed high enough through it to look innocent [2].
That contradiction is not a monitoring bug. A hit ratio describes what happened after a key was already warm; it says nothing about how many cold-start requests became database work in the same instant [4]. The service cached tenant project lists for thirty seconds [3], so the window in which a key is unowned and cold is roughly one query duration long. In the published repro the query is stubbed at 0.1 seconds [11], which makes the dangerous window about 0.33 percent of the TTL [15]. Aggregate ratios are averages over minutes. They cannot resolve a hazard that lives for a fraction of a second per key, several hundred times smaller than the refresh interval it hides inside.
The repro is worth running because it is honest about the mechanism. Twenty threads hit one tenant endpoint, and the counter prints db_calls: 20 [5]. The first request checked the cache, missed, released the lock, and went to the database; the other nineteen did exactly the same thing, because nothing had marked the key as being loaded [6]. The author names this a cache stampede [7]. Nineteen of the twenty queries were redundant, or 95 percent of the burst's database work [16].
Note what the author did before touching the code: he made the failure measurable, and set the bar that any change failing to move db_calls from twenty toward one was not solving the race [8]. That is the right order of operations, and it also tells you which metric belongs on the dashboard. Not hit ratio. Concurrent cold misses per key.
The fix supplies the instrument for free. The singleflight wrapper keeps an inflight dictionary mapping each key to an event and a result box; the first caller runs the loader and the rest wait on the event and receive the same result [9]. Rerun the burst and db_calls drops to one, so the database sees a single query while the cache remains useful afterwards [10]. The size of that inflight map, and the count of waiters per entry, is exactly the signal a hit ratio cannot express. Export it as a gauge. An alert on waiters-per-key above a small threshold fires during the stampede, not in the incident review afterwards.
Two caveats on the material. The author states he used MonkeyCode's free model access to compare a per-key lock against a process-wide lock before committing [12], and discloses that the article was prepared as part of MonkeyCode's product outreach [13]. Treat the tooling mention accordingly; the repro and the counter stand on their own. The author also warns that thread scheduling can hide a race [14], which is the reason a single passing burst test is weak evidence and a per-key concurrency gauge in production is strong evidence.
What to watch: whether your cache client exposes in-flight loads at all. Most expose hits, misses, and evictions, which means the number that predicts the outage is the one you have to add yourself. Check also whether your singleflight is per process. Twenty threads deduplicated to one query per replica is still one query per replica, and a fleet-wide release still lands as a burst on the database.