Build1 publisherNot yet confirmed elsewhere2 min readPublished
Speculative decoding in llama-server swaps real logprobs for 0.0 placeholders
llama-server b11430 reports logprob 0.0 for every speculatively decoded token, dragging one test's mean logprob from -0.48 to -0.0011. Nothing in the response or the server log flags the fill-ins, so evals and calibration built on those numbers go wrong quietly.
The Engineer · Build desk
What happened
- In a CPU reproduction, all 23 tokens after the first came back as logprob 0 with no alternatives, at temperature 0 and 1, on both the chat-completions and native endpoints.
- N-gram speculation drafts only when text repeats, so one response mixes real values and fill-ins: 3 of 30 tokens with default settings, 26 of 30 with a shorter match length.
- The request fields that once tuned speculation are compiled out and ignored, so a client cannot switch speculation off for a single request.
- Build b10703, from August, returned the same placeholders on the same probes, so the fault predates the current release.
- The sampled tokens are sound: on a small fixture, the server's speculative output matched plain sampling's next-token distribution.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure Scores built on these logprobs report near-certainty: implied per-token perplexity drops from about 1.62 to 1.001 on text the model sampled at temperature 1.
- constraint Filtering for zero logprobs cannot find the fill-ins because real tokens also score 0; the usable in-band signal is an empty top_logprobs on a request that asked for alternatives.
- contradiction A config audit that reads /props will clear a speculating server, since /props reports speculation as none while /slots reports it as on.
- cost Keeping speculative throughput and real probabilities together means running a second, unspeculated instance and routing every logprob request to it.
The fill-in starts in tools/server/server-context.cpp, which has two places that emit a generated token [22]. Both write a probability of 1.0 first [22][4]. The normal path does it under the comment "// TODO: set it here instead of doing inside populate_token_probs" [22]. The speculative path does it under "// set later", and according to the write-up nothing comes back to set it [4]. Stored as a logprob, a probability of 1.0 becomes ln(1) = 0 [25]. Later has not arrived by b11430 [4].
The reproduction ran CPU-only in a Debian 13 LXC container, on the b11430 release tarball with its SHA-256 matched to the release digest, with gemma-3-1b-it-Q4_K_M as the model [11]. A draft model normally has to be smaller than the target to save time [12]. Pointing -md at the same file sends every token through the speculative path and needs no second download [12]. The native probability field breaks the same way [14]. With post_sampling_probs set, the speculating server returned prob 1.0 and an empty top_probs after the first token, where the unspeculated server gave "What" a probability of 0.048 [14].
Across five seeds of 64 tokens at temperature 1, the unspeculated per-seed means ran from -0.3360 to -0.6775 [15]. With the draft model, four seeds came back as -0.0 and the fifth as -0.0055 [15]. The bug report the write-up reproduces saw the same pattern on a Vulkan iGPU running Gemma 4 26B with an MTP drafter: a mean of -0.00107 with MTP against -0.38244 without [23][16]. That reporter was trying to find out whether MTP changes the output [16].
Those baselines come from two Gemma models on particular prompts, and another model will give another baseline [15][16]. The structure carries over. With a draft model, every token after the first is a placeholder [3]. A sequence's summed logprob is then the first token's value alone, whatever the model or prompt [26].
A schema check passes. The fields carry the expected types, a numeric logprob and a list for top_logprobs, so the response clears any schema written for it [18]. The generation side holds up better than the project's own example: the llama-speculative example program failed the next-token distribution test that the server's speculative output passed [2][1].
The write-up's toolkit includes check-spec-logprobs.sh, which checks a running server [9]. In my view, a harness that scores logprobs should run a check like that before every eval and refuse to score when it fails. The fault is tracked upstream as ggml-org/llama.cpp #27972 and #29975 [10].
What to watch
- A fix under ggml-org/llama.cpp #27972 or #29975 that fills in real probabilities on the speculative path, or at least marks the placeholders in the response.
- A correction to /props so it reports the active speculation types, letting config audits find affected servers without querying /slots.
- A return of per-request speculation controls, letting one shared server answer logprob requests without drafting.