Build1 publisher3 min readPublished
NVIDIA's 4-bit Nemotron shows what aggressive quantization costs: you have to retrain for it
A 66 GB checkpoint becomes 22 GB and, NVIDIA says, up to 4x faster throughput, but only after distillation teaches the quantized student to live with its own noise.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The Nemotron 3.5 Lightning NVFP4 checkpoint is compressed down to 22 GB from the 66 GB full precision checkpoint by quantizing many of its weights to 4 bits.
- The Nemotron 3.5 Lightning NVFP4 checkpoint preserves accuracy while unlocking up to 4x faster throughput.
- Post-training quantization (PTQ) is a common method for compressing models to NVFP4 and covers most needs, but reaching high throughput with tighter memory requires more aggressive quantization, for which quantization-aware distillation (QAD) is described as an optimal choice.
- NVIDIA states that QAD recovers accuracy degradation from aggressive quantization, and that even with more conservative configurations QAD consistently outperforms PTQ on agentic benchmarks while reducing memory usage.
- The QAD process used for Nemotron 3.5 Lightning NVFP4 has two stages: Stage 1 is a PTQ pass that quantizes weights to W4A16 to produce the quantized student; Stage 2 trains the student with QAD using a distillation loss aligning it with the teacher, yielding the final NVFP4 checkpoint.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
NVIDIA has published the pipeline behind its Nemotron 3.5 Lightning NVFP4 checkpoint, which the company says is compressed to 22 GB from a 66 GB full-precision checkpoint by quantizing many of its weights to 4 bits, and which unlocks up to 4x faster throughput [1][2]. The part worth an operator's attention is the order of operations: the accuracy survives because the model was trained to absorb quantization noise, not because the rounding was clever [3][4].
NVIDIA puts post-training quantization in its place fairly bluntly. PTQ is the common method and covers most needs, but reaching high throughput with tighter memory requires more aggressive quantization, and for that the company points to quantization-aware distillation [3]. QAD is two stages [5]. Stage one runs PTQ on the full-precision BF16 base, NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, quantizing weights to W4A16 to produce a student; the same BF16 model is the frozen teacher [5][6]. Stage two trains that student with every forward pass going through simulated quantization, while a KL divergence loss over teacher and student logits pulls its behaviour back toward the original [7][8].
The most useful detail is not the architecture but the changed acceptance threshold. PTQ on its own normally targets a median accuracy recovery above 99%; when QAD is planned, NVIDIA says a 95-99% target is acceptable because QAD will recover more accuracy [9]. That is a decision to ship a deliberately worse stage-one artefact, roughly four points more median degradation at the low end [10], on the expectation that the second stage closes the gap. In practice it bought a specific structural choice: quantizing the Mamba linear layers to W4A16 rather than FP8, which NVIDIA reports unlocked higher throughput without a large accuracy drop [11].
The calibration choice is not free either. NVIDIA tried several PTQ recipes differing in how weights are calibrated and how aggressively the Mamba projections and KV cache are quantized, and says that choice carries forward into training, with max-calibrated recipes feeding dynamic-scale QAD [12]. So the stage-one recipe is not a disposable first draft; it commits you to a particular stage-two configuration.
Two pieces of arithmetic are worth doing yourself. 66 GB to 22 GB is a 3x reduction, short of the 4x you would get by moving every 16-bit weight to 4 bits, which is consistent with NVIDIA's own wording that many, not all, weights are quantized [13]. And the stage-two training loop needs the frozen full-precision model resident alongside the student [6][7], so the 22 GB serving footprint is bought with a training run that still has to carry the 66 GB teacher.
What to watch: the throughput and accuracy figures here are NVIDIA's own, published on its developer blog [14], and the claim that QAD consistently outperforms PTQ on agentic benchmarks even under more conservative configurations is stated without published scores in this material [4]. The interesting number for anyone considering this route is the one NVIDIA has not put beside the 4x, which is how many GPU-hours of distillation the recovery costs. Until that is public, treat 4-bit as a training project with a serving benefit, not a conversion step.