Build1 publisherNot yet confirmed elsewhere3 min readPublished
llama-server's default prompt cache mixes two LoRA scales into one answer
llama-server's default prompt cache restored KV built under the wrong LoRA scale on 10 of 10 test prompts in release b11514. According to the dev.to write-up that found it, servers that swap adapters per request behind a shared system prompt trigger it all day.
The Engineer · Build desk
What happened
- llama-server's LoRA check compares an incoming request with the slot's previous request, so it never looks at the cache entry it has just restored.
- The failing sequence sends a prompt at scale 0, then a different prompt that takes the only slot, then the first prompt again at scale 1.
- Changing the adapter scale globally through POST /lora-adapters produced the same wrong answers on all ten test prompts.
- Restarting with --cache-ram 0, or sending cache_prompt: false on the repeat request, produced the correct answer for every prompt.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Operators who switch adapters per request must give up RAM-cache prefix reuse to get answers computed wholly at the requested scale, and the throughput cost is theirs to measure.
- constraint A LoRA test that sends each setting once, or two settings back to back, passes on an affected server, so a regression test needs a slot-stealing request between the two settings.
- exposure On a production persona or style adapter a half-and-half answer looks only slightly off, so affected deployments are more likely to see quality drift than an error they can trace.
At step three, the only slot last served Q, also at scale 1. When P comes back at scale 1, the slot-level check compares two scale-1 requests and finds nothing to invalidate [3]. The RAM cache then restores the KV it saved for P in step one. That KV was computed at scale 0, and the entry does not record the scale [1][2]. For the first prompt the timings read cache_n=116 and prompt_n=1. So the cache supplied the entire prompt, and the model evaluated just a single token [5]. The 48 output tokens were then generated with scale 1 applied [11][19]. The author's explanation is that the prompt was read under scale 0 and the continuation under scale 1, and the output matched neither reference answer [6].
The unit the cache stores is struct server_prompt in tools/server/server-task.h. At b11514 it starts with the prompt's tokens and a list of prompt checkpoints [18]. A correct restore check needs the adapter list and scales stored in that entry, then compared against the incoming request [2][3]. Today the only comparison is against whatever the slot ran last [3].
Nothing in the response marks which tokens came from the cache. "cache_n tells you that something was reused, not under what," the author wrote [21]. The trigger is the same prompt prefix under two LoRA settings, with a request in between that takes the slot [22]. The author points out that a server switching adapters per request, with a shared system prompt and a persona or task per adapter, does exactly that all day [22]. With -np 4 the same ten-prompt sequence came out right, because each prompt kept its own slot [15]. I'd expect extra slots only to delay the bug. The restore path starts whenever another request takes a slot [1], and a busy server with more distinct prefixes than slots will keep doing that.
The test design is sound. Reference answers came from a separate server started with --cache-ram 0, where the two scales gave different answers on all ten prompts [12]. The setup was CPU-only in a Debian 13 container, on a 15M-parameter story model and its Shakespeare adapter. That is the pair llama.cpp's own server tests use for LoRA [10]. I think the 10 of 10 [4] transfers to larger models, because the faulty comparison is between request settings and never touches the weights [3]. What has to hold in production is the access pattern.
The control row was not perfectly clean. With step one also at scale 1, one of ten answers differed, the same prompt each time, and it matched on each of three solo re-runs [13]. The author attributes that to a floating-point difference from building the KV in a different order, and says the cause was not proven [13].
The August build b10703 failed on all ten prompts as well [8]. Upstream tracks the per-request case as ggml-org/llama.cpp#30129 and the global path as #26207 [16]. The author's check-lora-cache.sh tests a running server for the behaviour [17].
What to watch
- Whether the fix for ggml-org/llama.cpp#30129 stores adapter settings with the cache entry and compares them on restore, and which build carries it.
- Whether #26207, the POST /lora-adapters global path, is closed by the same change or needs its own.
- A reproduction on a production-size model with a persona adapter, showing how visible a mixed-scale answer is outside the 15M test model.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence68
- Adoption
- Insufficient
- Hype gap+10
- Incentives20
- Confidence62
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
llama-server's RAM prompt cache (--cache-ram, on by default) saves a slot's KV when another request takes the slot, and restores it when the same prompt comes back.
- [2]
The saved cache entry does not record which LoRA adapters and scales the KV was computed with.
- [3]
The server checks for a LoRA change, but it compares the new request with the slot's previous request, not with the cache entry it has just restored.
- [4]
On b11514, the current release, the sequence decoded the repeat request on scale-0 KV for 10 of 10 prompts; the answer differed from the same request without the cache.
- [5]
For the first prompt, step three timings showed cache_n=116 and prompt_n=1: the whole prompt came from the cache and one token was evaluated.
- [6]
The step-three answer was neither the scale-1 answer nor the scale-0 one: the prompt was read under scale 0 and the continuation generated under scale 1.
- [7]
Changing the scale globally with POST /lora-adapters instead produced answers differing from the scale-1 reference on 10 of 10 prompts.
- [8]
Build b10703 from August, at defaults, differed from the scale-1 answer on 10 of 10 prompts.
- [9]
With --cache-ram 0, or with step three sent with cache_prompt: false, 0 of 10 answers differed from the scale-1 reference.
- [10]
The test ran in a Debian 13 LXC, CPU only, with llama.cpp b11514 from the release tarball, using stories15M_MOE-F16.gguf (a 15M-parameter story model) and its moe_shakespeare15M.gguf adapter, the models llama.cpp's own server tests use for LoRA.
- [11]
The server ran with --lora-init-without-apply -np 1 -c 4096; ten prompts of about 116 tokens each were sent at temperature 0 with 48 tokens out.
- [12]
Reference answers came from a separate server started with --cache-ram 0; the two scales gave different answers on all ten prompts.
- [13]
With step one also at scale 1, 1 of 10 answers differed, the same prompt each time; re-run on its own three times it matched. The author says a floating-point difference from building the KV in a different order is the likely cause, but that was not proven.
- [14]
A test that sends one request per setting, or the two settings back to back, never sees the bug, because the slot still holds the KV and the slot-level check works.
- [15]
With -np 4 and the same ten-prompt sequence, every answer was right because each prompt kept its own slot.
- [16]
Upstream issues are ggml-org/llama.cpp#30129, and #26207 for the global path.
- [17]
The author's toolkit script check-lora-cache.sh tests a running server.
- [18]
The unit the cache stores, in tools/server/server-task.h at b11514, is struct server_prompt, which begins with server_tokens tokens and a std::list of common_prompt_checkpoint.
- [19]
The generated tokens are produced with the requested adapter applied; only the part of the context that came from the cache was computed differently.
- [20]
The reproduction sends prompt P with adapter 0 at scale 0, then prompt Q at scale 1 which takes the only slot, then P again at scale 1.
- [21]
cache_n tells you that something was reused, not under what.
ReportedInsufficientSource: dev.to write-up author2 sources— create a free account to open themView cited source - [22]
The bug needs the same prompt prefix sent under two different LoRA settings with another request in between that takes the slot; the author says that is what a server switching adapters per request does all day, with a shared system prompt and different personas or tasks per adapter.
ReportedInsufficientSource: dev.to write-up2 sources— create a free account to open themView cited source - [23]
On a real model with a style or persona adapter, a request that is half one adapter and half the other looks like a slightly off answer.
ReportedInsufficientSource: dev.to write-up2 sources— create a free account to open themView cited source
Sources
1 independent publisher whose own reporting we read for this story.
- dev.tollama-server's prompt cache reuses KV computed under a different LoRA scale
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.