Repacking Google's quantization-aware-trained Gemma 4 weights without re-rounding fits the 12B model on one TPU v5e chip at bf16-level accuracy. Adopting it means carrying patches to vLLM's TPU backend and working inside a 9,728-token KV cache.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence55
Repacking Gemma 4's QAT weights with 4-bit embedding tables made decode up to 1.39x faster on one SageMaker L4, according to a dev.to benchmark series. For teams serving Gemma 4 on vLLM, how the weights are stored becomes a setting to measure alongside model size.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
Quantized to about 4 GB, a 7B model still runs out of GPU memory near 30,000 tokens as its KV cache grows, according to a dev.to walkthrough. Sizing a card by its weight file alone leaves that growing cache off the memory budget.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+15
- Incentives25
- Confidence40
Absmax INT8 takes one scale from the largest magnitude in a tensor, so a single outlier sets the step size for everything else. The LLM.int8() authors found those outliers sitting in about six feature dimensions.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence58
NVIDIA's submission runs the same Qwen3.6-27B as the llama.cpp reference on the same Jetson board and finishes 6.4x sooner. Most of the gap comes from prompt tokens the runtime never has to prefill.
Reality
- Evidence58
- Adoption18
- Hype gap+20
- Incentives85
- Confidence62
An August 2026 paper argues low-bit quantization-aware training converges high because its reconstruction step ignores which weights matter. The fix is reported to cost 1.4% of step time.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+34
- Incentives
- Insufficient
- Confidence27