Skip to content

Build1 publisher3 min readPublished

Speculative decoding is free because the bottleneck is the memory bus, not the math

A conversion-pipeline checklist item, "MTP round-trip", turns on a distinction teams collapse: the training-time auxiliary loss is disposable, the inference-time draft head is not.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The author saw "MTP round-trip" listed as an item to verify on a checklist for a Megatron conversion pipeline, and did not initially know what it meant.
  • To produce token N+1 the model needs token N; there is no way around that ordering, which is what "language model" means.
  • Generating 100 tokens means 100 full forward passes through the network.
  • The post reports GLM-5.2's config as "num_hidden_layers": 78 and "num_nextn_predict_layers": 1.
  • A forward pass processing one token and a forward pass processing five tokens take roughly the same wall-clock time.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A checklist line on a Megatron conversion pipeline reading "MTP round-trip" is the kind of item that gets ticked without being understood [1]. It is worth understanding, because what it verifies is the difference between a model that emits one token per weight read and one that emits several, at identical output quality [8][10].

Start with the constraint everyone knows. To produce token N+1 the model needs token N, and that ordering is what "language model" means [2]. Generating 100 tokens therefore means 100 full forward passes [3]. For a model with 78 hidden layers, which is what the post reports GLM-5.2's config declaring, that is 7,800 layer traversals to write a short paragraph [4][1].

The intuition that fails next is the pricing. A forward pass over one token and a forward pass over five take roughly the same wall-clock time [5]. Every pass has to pull the model's weights out of memory into the compute units, hundreds of gigabytes across the bus, whether the batch is one token or fifty [6]. The arithmetic on a handful of tokens is rounding error against the cost of fetching the weights to do it with [6]. Same tokens, same math, five times the memory traffic, purely because of the ordering [7]. The ratio is not a subtlety, it is the entire opportunity.

Speculative decoding collects it by splitting generation into two roles: something cheap drafts k tokens, the large model does one pass over all k positions and checks them, the longest correct prefix is accepted and the rest discarded [8]. Verification parallelises even though generation does not, because the candidate tokens already exist when verification starts [9]. Three correct guesses yield four tokens for the price of one weight read [11], four times the output per unit of memory traffic [2]. A wrong guess on the first position yields one token, which is the baseline [12].

The part that usually hides a catch does not. Per the post, the acceptance test is constructed so surviving tokens are distributed exactly as if the large model had generated them alone, so the output distribution is unchanged [10]. There is no quality knob and no regression to monitor [13]. The only cost of a bad drafter is wasted drafting work [14], which makes acceptance rate the sole metric that matters [15].

That is why drafter provenance matters. A separately trained draft model does not necessarily think like the target, acceptance falls, and a drafter whose guesses are rejected is pure overhead [16]. Building the drafter into the model is the response, and Multi-Token Prediction is that [17]. The config the post quotes shows the asymmetry: 78 hidden layers, one next-token prediction layer [4], roughly 1.3 percent of the layer count [3], though layer count is not parameter count.

Here is where pipelines get bitten. MTP has two lives: a training-time auxiliary loss you can throw away, and an inference-time draft head you cannot [18]. Same acronym, different object. The post does not describe a failure, but the structure implies one: because drafts never alter the output distribution [10], a checkpoint that loses its prediction layer in conversion still produces correct text and simply runs slower [4]. That failure passes every correctness test a team is likely to own.

Two things to check. Whether your converted checkpoint still declares a next-token prediction layer after a round trip [1][4], and whether you are logging acceptance rate per workload rather than assuming the drafter earns its keep [15].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories