Liquid AI shipped a 279.5-million-parameter draft model for LFM2.5-VL-3B on September 24, reporting 2.30x to 3.13x faster decoding on an Apple M5 Max. Image encoding and prompt prefill are unchanged, and the company's tests did not cover quantized deployments.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence60
Inferact measured 709 output tokens a second on 16 Ironwood chips against 452 on 16 GB200s, at low concurrency with speculative decoding. The engineering worth reading is the hand-written memory schedule underneath.
Reality
- Evidence45
- Adoption22
- Hype gap+20
- Incentives78
- Confidence58
DFlash uses a diffusion model to draft three tokens at once and claims to beat EAGLE-3. On one consumer GPU running llama.cpp on a JavaScript coding task, it trailed the drafter Google ships with Gemma-4-12B-it.
Reality
- Evidence45
- Adoption25
- Hype gap+30
- Incentives35
- Confidence50
The release attaches speculative decoding to weights it publishes under MIT, which lowers what a self-hosting threat costs to stand up, even though the performance claim behind it is still the vendor's own.
Reality
- Evidence46
- Adoption38
- Hype gap+32
- Incentives78
- Confidence55
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58
A week-long failure log on two RTX 3090s under WSL2 lands on one config at 170-210 tok/s. Everything before it died in dependency resolution, not in the math.
Reality
- Evidence38
- Adoption18
- Hype gap−12
- Incentives27
- Confidence44