Speculative decoding in llama-server swaps real logprobs for 0.0 placeholders
llama-server b11430 reports logprob 0.0 for every speculatively decoded token, dragging one test's mean logprob from -0.48 to -0.0011. Nothing in the response or the server log flags the fill-ins, so evals and calibration built on those numbers go wrong quietly.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence58