KAIST and Seoul National University's AgSpec nearly doubles accepted draft length for coding agents by indexing files in the diff and JSON forms agents emit. It lives entirely in the retrieval index and leaves model weights untouched.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
g factor's Qwen 3.8 27B benchmark has Together AI fastest at one stream, at 189.61 tok/s, while four of five engines finish within about 10% at 64 streams. Choosing a provider from these numbers starts with knowing how many streams the deployment will run at once.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+18
- Incentives72
- Confidence45
Liquid AI shipped a 279.5-million-parameter draft model for LFM2.5-VL-3B on September 24, reporting 2.30x to 3.13x faster decoding on an Apple M5 Max. Image encoding and prompt prefill are unchanged, and the company's tests did not cover quantized deployments.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence60
Inferact measured 709 output tokens a second on 16 Ironwood chips against 452 on 16 GB200s, at low concurrency with speculative decoding. The engineering worth reading is the hand-written memory schedule underneath.
Reality
- Evidence45
- Adoption22
- Hype gap+20
- Incentives78
- Confidence58
AWS counted 13 SageMaker inference launches so far in 2026. The one that changes production behaviour most is a prioritized list of up to five instance types, each allowed its own model optimization settings.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives85
- Confidence55
NVIDIA's submission runs the same Qwen3.6-27B as the llama.cpp reference on the same Jetson board and finishes 6.4x sooner. Most of the gap comes from prompt tokens the runtime never has to prefill.
Reality
- Evidence58
- Adoption18
- Hype gap+20
- Incentives85
- Confidence62
A September 3 playbook traces speculative decoding's draft architectures from EAGLE-3 to DFlash, and grounds the case in a 70B model that decodes at 15 to 20 tokens a second on eight H100s.
Reality
- Evidence24
- Adoption31
- Hype gap+38
- Incentives38
- Confidence33
Alibaba's 27B scores 52 on the Artificial Analysis index from a 17GB quantized file. Filling its 262,144-token window needs roughly 16 GiB of KV cache on top of that, so the file size is the smaller half of the sizing question.
Reality
- Evidence30
- Adoption15
- Hype gap+45
- Incentives50
- Confidence32
Red Hat says 70 to 80 percent of enterprise AI spending goes to inference, and blames a 70-billion-parameter model reading all 140GB of its weights for every token it writes. Its fix needs a draft model you train yourself.
Reality
- Evidence55
- Adoption40
- Hype gap+20
- Incentives85
- Confidence45
DFlash uses a diffusion model to draft three tokens at once and claims to beat EAGLE-3. On one consumer GPU running llama.cpp on a JavaScript coding task, it trailed the drafter Google ships with Gemma-4-12B-it.
Reality
- Evidence45
- Adoption25
- Hype gap+30
- Incentives35
- Confidence50
NVIDIA reports 2.5x more concurrent users for Nemotron 3 Ultra from its packaged NIM serving stack. The gain comes from five interacting layers measured on one agentic traffic shape, and NVIDIA ships the harness to retest it.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+18
- Incentives85
- Confidence62
The release attaches speculative decoding to weights it publishes under MIT, which lowers what a self-hosting threat costs to stand up, even though the performance claim behind it is still the vendor's own.
Reality
- Evidence46
- Adoption38
- Hype gap+32
- Incentives78
- Confidence55
A new paper fits log-linear laws for draft acceptance rate against pretraining tokens, draft capacity and decoding batch size, and argues with a roofline model that the tree you tune at batch 1 is the wrong tree in production.
Reality
- Evidence38
- Adoption8
- Hype gap+32
- Incentives70
- Confidence46
Speculative decoding speedups depend on the data, and most published ones come from high-level scripts on narrow datasets. SPEED-Bench's authors argue the honest measurement happens inside vLLM or TensorRT-LLM, across concurrencies.
Reality
- Evidence42
- Adoption21
- Hype gap+18
- Incentives58
- Confidence44
A 7B model split across Iowa and Oregon on free T4s went from 4.92 to 28.10 tokens per second. Most of the gain came from a drafter that stopped launching kernels one at a time.
Reality
- Evidence42
- Adoption16
- Hype gap+22
- Incentives58
- Confidence38
A published NVFP4 and speculative-decoding config turns a 27B open-weights model into something you can try to serve. The 206.1 tokens per second figure is single-stream and unreplicated.
Reality
- Evidence52
- Adoption28
- Hype gap+14
- Incentives70
- Confidence46
A conversion-pipeline checklist item, "MTP round-trip", turns on a distinction teams collapse: the training-time auxiliary loss is disposable, the inference-time draft head is not.
Reality
- Evidence42
- Adoption55
- Hype gap+22
- Incentives18
- Confidence45
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58
A week-long failure log on two RTX 3090s under WSL2 lands on one config at 170-210 tok/s. Everything before it died in dependency resolution, not in the math.
Reality
- Evidence38
- Adoption18
- Hype gap−12
- Incentives27
- Confidence44