Skip to content

Build1 publisher3 min readPublished

Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters

The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Inco AI released DFlash 2, an updated version of the parallel speculative-decoding method DFlash, on August 18.
  • Inco AI reports a 21% increase in average acceptance length over the original DFlash.
  • The path selector and convolution responsible for the reported gain added 1.3% to draft-verify cycle latency in Inco AI's tests.
  • Inco AI's published Qwen benchmark uses a five-layer Qwen3-4B DFlash model on GSM8K.
  • Inco AI says the final output remains unchanged because the larger target model still verifies every proposed token.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Inco AI released DFlash 2 on August 18, an update to its parallel speculative-decoding method, and reports a 21% increase in average acceptance length for a 1.3% increase in draft-verify cycle latency [1][2][3]. The number worth arguing about is not the ratio but where it comes from: 18.5 million parameters of scaffolding bolted onto an existing five-layer drafter, rather than a bigger model [1][12][16].

Speculative decoding has a small model propose several future tokens, then has the expensive target model check them together; accepted guesses save forward passes, rejected ones are discarded, and the target model's output is preserved [9]. DFlash's contribution, in a paper submitted to arXiv in February by UC San Diego-affiliated authors Zhijian Liu, Jian Chen and Yesheng Liang, was to drop the sequential drafting loop and propose every position in a single parallel pass [6][7][8]. That trade has a known failure mode: tokens that are individually plausible can read badly when placed next to one another [10].

DFlash 2 answers with two cheap parts. A path selector keeps the top 16 candidates at each position, scores neighboring token pairs and walks those precomputed scores to pick a coherent sequence, at a cost of 2 million parameters and 0.6% cycle latency [11][12]. In Inco AI's Qwen3-4B test it produced longer accepted drafts than a DSpark correction module while using roughly 40 times fewer parameters [13]. The second part targets what Inco AI calls suffix decay, the drafter's falling accuracy near the end of a proposed block [14]. A two-tap local convolution around each attention and feed-forward sublayer lets every position mix with its immediate predecessor while keeping the computation parallel, for 16.5 million parameters, about 3% of the drafter, and 0.7% cycle latency [15][16].

The instructive comparison is the alternative Inco AI measured. Tripling the drafter from five layers to 15 recovered much the same end-of-block accuracy at a 15.2% latency cost, roughly 20 times the convolution's [17][3]. Similar accuracy, very different bill. Gains across the benchmark suite ranged from 16% to 25%, with the published Qwen figure from a five-layer Qwen3-4B drafter on GSM8K [18][4]. Output does not change, because the target model still verifies every proposed token [5].

These are the vendor's own numbers, and Inco AI concedes that headline throughput depends on model, task, sampling settings, hardware and concurrency [19]. It reports 2.7x to 3.4x autoregressive throughput for a Qwen3.8-27B drafter in SGLang at batch size one, and 3.1x to 4.6x for its Muse Glimmer drafter [20][21]. Independent results exist for the first DFlash, not DFlash 2: NVIDIA reported up to 15x higher throughput for gpt-oss-120b on an eight-GPU Blackwell system at the same interactivity target, and Google reported a 3.13x average speedup on TPU v5p in a standalone JAX benchmark plus a 2.29x end-to-end gain in a vLLM TPU pipeline comparison [24][25][26]. Those validate the original architecture across accelerators [27].

For anyone running a fleet, distribution matters more than the acceptance curve. The original repository is supported in SGLang, vLLM, TensorRT-LLM and llama.cpp, and the DFlash 2 release ships configurations for SGLang, vLLM, llama.cpp and an oMLX build for Apple Silicon [22][23]. Inco AI positions DFlash 2 as the first piece of an end-to-end inference stack for agent workloads, where a single task may run for hours and generate far more tokens than a chat session [28][29]. Agent traffic multiplies token demand and serving cost, and the pitch is fewer target-model passes without a change in output [30].

What to watch: acceptance-length measurements from someone other than Inco AI, at batch sizes above one, since the throughput figures are reported at batch size one [20][19]. Also whether the 16% to 25% spread narrows or widens on non-GSM8K work [18][4], and whether third-party accelerator tests move from the original DFlash to DFlash 2 [24][27].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories