A September 3 playbook traces speculative decoding's draft architectures from EAGLE-3 to DFlash, and grounds the case in a 70B model that decodes at 15 to 20 tokens a second on eight H100s.
Reality
- Evidence24
- Adoption31
- Hype gap+38
- Incentives38
- Confidence33
Retrieve-for-Train runs reinforcement learning once against a fixed corpus and distils the winning sub-query sets into a lightweight diffusion model. Adopting it puts a retrain schedule on whoever owns the catalogue.
Reality
- Evidence45
- Adoption8
- Hype gap+30
- Incentives60
- Confidence55
Across 20 paired writing-correction cases on Ollama for Windows, the two models succeeded on the same 18 and failed on the same 2, while the 4B averaged 23.99 seconds of cold start against 54.37. That reorders local shortlisting.
Reality
- Evidence46
- Adoption14
- Hype gap+16
- Incentives42
- Confidence55
Perturbation probing needs the weights, which puts the finding on teams running open-source checkpoints rather than on API tenants. The same 50-neuron toolkit that breaks a refusal template also repairs other behavior.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+30
- Incentives80
- Confidence45
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58