Build1 publisher3 min readPublished
Your LLM cache probably never fires, because it hashes spelling instead of meaning
A dev.to writeup traces a support bot's dead cache to hash-based lookup and replaces it with cosine similarity over embeddings at a 0.92 threshold. The lever sits in retrieval, not model choice.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A post titled "Semantic Cache in AI Tokenomics" on dev.to describes the author reviewing logs for a support bot that answers the same handful of questions all day long.
- The logged question variants cited were: "How do I reset my password", "how do I reset my login", "I forgot my password, help", and "can't log in, need to reset it" - four different sentences with one question underneath them.
- The bot had a cache but it never hit; every one of those four questions triggered a full model call because the cache was checking for an exact string match and none of the sentences matched character for character.
- A normal cache works off a hash of the input: same input, same hash, same cached response. The post calls this fast, simple, and effective for a huge class of problems.
- The post argues exact-match caching falls apart when the input is natural language written by different humans, giving the example that "What are your hours" and "when are you open" mean the same thing to a person and mean nothing alike to a hash function.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A dev.to post titled "Semantic Cache in AI Tokenomics" opens with an unglamorous log review: a support bot answering the same handful of questions all day, sitting behind a cache that never once hit [1][3]. It matters because the money was leaking in the lookup layer, not in the model, and that is a layer most teams never instrument.
The failure is mechanical. A conventional cache keys on a hash of the input: same input, same hash, same stored response, which is fast, simple, and correct for a large class of problems [4]. Natural language typed by different humans is not that class. The author's logged variants were "How do I reset my password", "how do I reset my login", "I forgot my password, help", and "can't log in, need to reset it" [2] - one question underneath four strings, and therefore four hashes and four full model calls [3]. As the post puts it, "What are your hours" and "when are you open" mean the same thing to a person and nothing alike to a hash function [5]. The author's argument is that the fix is not a bigger cache but a different lookup: meaning instead of spelling [6].
The construction has three parts. An embedding turns a sentence into a list of numbers via a call to an embedding model, which most providers offer, such that similar meanings land on similar numbers regardless of wording [7]. A similarity score between 0 and 1, computed with cosine similarity, says how close two embeddings are [8]. A threshold, in the post's example 0.92, decides how close is close enough to reuse a stored answer rather than pay for a new call [9]. The accompanying code is four lines of numpy for the similarity function and a cache class that keeps a list of prior answers and scans new questions against it, though the excerpt cuts off mid-class [11].
The worked example is the useful part. With one entry stored for "How do I reset my password", the two rephrasings of it match despite barely overlapping as text, while a genuinely different third question lands far away and correctly triggers a real model call [10]. That is two of three incoming questions served from cache in the illustration, roughly a 67 percent hit rate against zero for the hash [13]. Treat that as a diagram, not a benchmark.
The threshold is where the engineering judgment sits, and it cuts both ways: the same number that lets "reset my login" reuse a password answer will, if set too loosely, hand that answer to a question that only sounds adjacent [14]. Note also what the piece does not supply: no measured hit rates, no latency figures, no cost for the per-lookup embedding call that a semantic cache adds to every request, and no treatment of invalidation [15]. The embedding call is not free, so the arithmetic only closes if hits are frequent enough to cover it.
The cheap first move is the audit the author suggests: pull a week of prompts from a customer-facing feature and group them by rough topic rather than exact text, and look for clusters that never hit each other [12]. Watch two numbers after that, not one. Hit rate tells you whether the cache is earning its embedding calls; a sampled review of what it returned tells you whether the threshold is quietly serving wrong answers faster than before.