Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

llama-server's default prompt cache mixes two LoRA scales into one answer

llama-server's default prompt cache restored KV built under the wrong LoRA scale on 10 of 10 test prompts in release b11514. According to the dev.to write-up that found it, servers that swap adapters per request behind a shared system prompt trigger it all day.

The Engineer · Build desk

How we use AISend a correction

What happened

  • llama-server's LoRA check compares an incoming request with the slot's previous request, so it never looks at the cache entry it has just restored.
  • The failing sequence sends a prompt at scale 0, then a different prompt that takes the only slot, then the first prompt again at scale 1.
  • Changing the adapter scale globally through POST /lora-adapters produced the same wrong answers on all ten test prompts.
  • Restarting with --cache-ram 0, or sending cache_prompt: false on the repeat request, produced the correct answer for every prompt.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Operators who switch adapters per request must give up RAM-cache prefix reuse to get answers computed wholly at the requested scale, and the throughput cost is theirs to measure.
  • constraint A LoRA test that sends each setting once, or two settings back to back, passes on an affected server, so a regression test needs a slot-stealing request between the two settings.
  • exposure On a production persona or style adapter a half-and-half answer looks only slightly off, so affected deployments are more likely to see quality drift than an error they can trace.

At step three, the only slot last served Q, also at scale 1. When P comes back at scale 1, the slot-level check compares two scale-1 requests and finds nothing to invalidate [3]. The RAM cache then restores the KV it saved for P in step one. That KV was computed at scale 0, and the entry does not record the scale [1][2]. For the first prompt the timings read cache_n=116 and prompt_n=1. So the cache supplied the entire prompt, and the model evaluated just a single token [5]. The 48 output tokens were then generated with scale 1 applied [11][19]. The author's explanation is that the prompt was read under scale 0 and the continuation under scale 1, and the output matched neither reference answer [6].

The unit the cache stores is struct server_prompt in tools/server/server-task.h. At b11514 it starts with the prompt's tokens and a list of prompt checkpoints [18]. A correct restore check needs the adapter list and scales stored in that entry, then compared against the incoming request [2][3]. Today the only comparison is against whatever the slot ran last [3].

Nothing in the response marks which tokens came from the cache. "cache_n tells you that something was reused, not under what," the author wrote [21]. The trigger is the same prompt prefix under two LoRA settings, with a request in between that takes the slot [22]. The author points out that a server switching adapters per request, with a shared system prompt and a persona or task per adapter, does exactly that all day [22]. With -np 4 the same ten-prompt sequence came out right, because each prompt kept its own slot [15]. I'd expect extra slots only to delay the bug. The restore path starts whenever another request takes a slot [1], and a busy server with more distinct prefixes than slots will keep doing that.

The test design is sound. Reference answers came from a separate server started with --cache-ram 0, where the two scales gave different answers on all ten prompts [12]. The setup was CPU-only in a Debian 13 container, on a 15M-parameter story model and its Shakespeare adapter. That is the pair llama.cpp's own server tests use for LoRA [10]. I think the 10 of 10 [4] transfers to larger models, because the faulty comparison is between request settings and never touches the weights [3]. What has to hold in production is the access pattern.

The control row was not perfectly clean. With step one also at scale 1, one of ten answers differed, the same prompt each time, and it matched on each of three solo re-runs [13]. The author attributes that to a floating-point difference from building the KV in a different order, and says the cause was not proven [13].

The August build b10703 failed on all ten prompts as well [8]. Upstream tracks the per-request case as ggml-org/llama.cpp#30129 and the global path as #26207 [16]. The author's check-lora-cache.sh tests a running server for the behaviour [17].

What to watch

  • Whether the fix for ggml-org/llama.cpp#30129 stores adapter settings with the cache entry and compares them on restore, and which build carries it.
  • Whether #26207, the POST /lora-adapters global path, is closed by the same change or needs its own.
  • A reproduction on a production-size model with a persona adapter, showing how visible a mixed-scale answer is outside the 15M test model.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence68
Adoption
Insufficient
Hype gap+10
Incentives20
Confidence62
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    llama-server's RAM prompt cache (--cache-ram, on by default) saves a slot's KV when another request takes the slot, and restores it when the same prompt comes back.

    ReportedSupportedSource: dev.to write-upView cited source
  2. [2]

    The saved cache entry does not record which LoRA adapters and scales the KV was computed with.

    ReportedSupportedSource: dev.to write-upView cited source
  3. [3]

    The server checks for a LoRA change, but it compares the new request with the slot's previous request, not with the cache entry it has just restored.

    ReportedSupportedSource: dev.to write-upView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 8, 2026

    llama-server's prompt cache reuses KV computed under a different LoRA scale

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories