Inferact measured 709 output tokens a second on 16 Ironwood chips against 452 on 16 GB200s, at low concurrency with speculative decoding. The engineering worth reading is the hand-written memory schedule underneath.
Reality
- Evidence45
- Adoption22
- Hype gap+20
- Incentives78
- Confidence58
A September 3 playbook traces speculative decoding's draft architectures from EAGLE-3 to DFlash, and grounds the case in a 70B model that decodes at 15 to 20 tokens a second on eight H100s.
Reality
- Evidence24
- Adoption31
- Hype gap+38
- Incentives38
- Confidence33
Decode is memory-bandwidth bound. Eight concurrent 128k Llama 3 70B streams need 97.8 ms of HBM transfer for every token generated, and DeepSeek's 512-scalar latent is aimed squarely at those bytes.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+40
- Incentives50
- Confidence50
The framework separates configuration choices that move a deployment along a latency and throughput frontier from advances that move the frontier itself, which is the distinction you need before trusting anyone's benchmark table.
Reality
- Evidence36
- Adoption24
- Hype gap+14
- Incentives81
- Confidence51
A conversion-pipeline checklist item, "MTP round-trip", turns on a distinction teams collapse: the training-time auxiliary loss is disposable, the inference-time draft head is not.
Reality
- Evidence42
- Adoption55
- Hype gap+22
- Incentives18
- Confidence45
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58
The French startup's only public number is 3,000 tokens per second on a 2-billion-parameter model. Its 30x claim for real LLMs has not been shown yet.
Reality
- Evidence30
- Adoption18
- Hype gap+52
- Incentives78
- Confidence38