build1 publisher
A diffusion drafter lost to Gemma's own Assistant model on a 12GB RTX 3060
DFlash uses a diffusion model to draft three tokens at once and claims to beat EAGLE-3. On one consumer GPU running llama.cpp on a JavaScript coding task, it trailed the drafter Google ships with Gemma-4-12B-it.
Publishers:dev.to
Reality
- Evidence45
- Adoption25
- Hype gap+30
- Incentives35
- Confidence50