Kimi K3's 1.4 terabytes of weights take eight Nvidia GB300s just to sit in memory. Export controls keep those chips away from Moonshot. Modal, Fireworks and Baseten price the hosted result at $3 in and $15 out.
Reality
- Evidence32
- Adoption45
- Hype gap+33
- Incentives74
- Confidence40
A September 3 playbook traces speculative decoding's draft architectures from EAGLE-3 to DFlash, and grounds the case in a 70B model that decodes at 15 to 20 tokens a second on eight H100s.
Reality
- Evidence24
- Adoption31
- Hype gap+38
- Incentives38
- Confidence33
DFlash uses a diffusion model to draft three tokens at once and claims to beat EAGLE-3. On one consumer GPU running llama.cpp on a JavaScript coding task, it trailed the drafter Google ships with Gemma-4-12B-it.
Reality
- Evidence45
- Adoption25
- Hype gap+30
- Incentives35
- Confidence50
Meta's Muse Glimmer 30B and Alibaba's Qwen3.8-27B both landed in August under pure Apache 2.0 and both fit one 24 GB GPU, so the deployment question moves off licence terms and onto how you spend the memory that is left.
Reality
- Evidence16
- Adoption12
- Hype gap+52
- Incentives62
- Confidence20
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58