build1 distinct publisher
SPEED-Bench re-tests speculative decoding at the batch size you actually serve
Speculative decoding speedups depend on the data, and most published ones come from high-level scripts on narrow datasets. SPEED-Bench's authors argue the honest measurement happens inside vLLM or TensorRT-LLM, across concurrencies.
Publishers:arxiv.org
Reality
- Evidence42
- Adoption21
- Hype gap+18
- Incentives58