Skip to content

Build1 publisher3 min readPublished Updated

QUASAR says the 2-bit QAT loss floor is a weighting bug, not a law of physics

An August 2026 paper argues low-bit quantization-aware training converges high because its reconstruction step ignores which weights matter. The fix is reported to cost 1.4% of step time.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Shrinking an LLM from 16-bit to 4-bit or 2-bit precision can cut memory requirements by 4-8x, making it possible to run large models on consumer hardware, edge devices, or cost-constrained cloud instances.
  • A paper from August 2026 titled 'QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction' identifies a specific, fixable cause of accuracy loss in Quantization-Aware Training and proposes a lightweight solution.
  • QUASAR reduces held-out KL divergence by up to 29% at 2-bit precision.
  • QUASAR adds only a 1.4% increase in training step time.
  • Post-training quantization methods such as GPTQ and AWQ quantize a fully trained model without further gradient updates; they are fast and require no training infrastructure, but struggle at very low bit-widths (2-bit, 3-bit) where rounding errors compound and accuracy drops sharply.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A write-up published on dev.to summarises an August 2026 paper, QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction, which locates the accuracy floor of low-bit QAT in a specific mechanical choice rather than in the arithmetic itself [2]. The reported payoff is up to 29% lower held-out KL divergence at 2-bit for a 1.4% increase in training step time [3][4], a ratio that makes this a cheap line item in an existing pipeline rather than a research programme.

The diagnosis is worth restating because it is not exotic. In standard QAT the forward pass runs on reconstructed weights, the quantized-then-dequantized tensor, while the optimizer updates the latent full-precision weights underneath [7]. Gradients are taken with respect to the reconstruction via the straight-through estimator and then applied to the latent weights, and because the two are linked by a lossy function, that signal is not the best descent direction for what is actually being updated [7]. The model still converges; it converges higher than a full-precision run, and the authors call that residual the loss floor gap [6][7].

QUASAR's theoretical claim is narrow and therefore useful: the saliency-weighted distance between the reconstruction and the ideal reconstruction is the only reconstruction-dependent term in the QAT convergence bound [8]. If that holds, improving the reconstruction is the whole lever. The implementation is three parts, all of them cheap. Per-parameter importance is approximated by an exponential moving average of squared gradients instead of a Hessian, using quantities the training step has already computed [9]. The affine dequantizer, scale and zero point, is then fitted in closed form to minimise saliency-weighted reconstruction error, so weights that move the loss get priority [10]. Finally a small candidate set of clipping ranges is searched, picking the one with the lowest weighted error before the dequantizer is fitted [11].

On cost: a 1.4% step-time increase turns a nominal 100-hour training run into 101.4 hours [13]. Compared against the headline 29% KL improvement, the benefit is roughly twenty times the cost in percentage terms, with the caveat that the two percentages measure unrelated quantities [14]. The context is that post-training methods such as GPTQ and AWQ need no training infrastructure at all but degrade sharply at 2-bit and 3-bit, which is exactly where QAT earns its keep [5][6], and where 16-bit to 2-bit compression is an eightfold reduction in weight memory [1][15].

Now the limits of what has actually been supplied. This is a secondary write-up, not the paper, and its text breaks off at the point where the 2-bit results are enumerated, so the per-model numbers behind the 29% figure are not in front of us [16]. The evaluation is described as covering the Qwen3 and Llama-3.1 families at 2, 3 and 4 bits against QAT and PTQ baselines [12], but "up to 29%" is doing unspecified work across that grid, and KL divergence against the full-precision teacher is not the same as downstream task accuracy.

What to watch: whether the 1.4% overhead holds at production tensor-parallel scale, since the closed-form fit and range search run per weight tensor per step [10][11]; and whether the improvement survives translation into task metrics rather than KL.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories