Repacking Google's quantization-aware-trained Gemma 4 weights without re-rounding fits the 12B model on one TPU v5e chip at bf16-level accuracy. Adopting it means carrying patches to vLLM's TPU backend and working inside a 9,728-token KV cache.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence55
Developer xbill9 rebuilt Google's QAT Gemma 4 26B as int4 and fit it on one TPU v6e with 53,888 tokens of KV cache, against 3,456 for RedHat's FP8 build. Throughput nearly doubles, with accuracy checked on one classification suite.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
Gemma 4 E2B's 4-bit QAT checkpoint decodes 2.05x faster than bf16 on one SageMaker L4, according to a dev.to benchmark. The swap also frees 18% of GPU memory for a 20% larger KV cache, though it runs only through the vLLM container because JumpStart lists no QAT build.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+8
- Incentives
- Insufficient
- Confidence55
A 66 GB checkpoint becomes 22 GB and, NVIDIA says, up to 4x faster throughput, but only after distillation teaches the quantized student to live with its own noise.
Reality
- Evidence38
- Adoption20
- Hype gap+32
- Incentives85
- Confidence46
An August 2026 paper argues low-bit quantization-aware training converges high because its reconstruction step ignores which weights matter. The fix is reported to cost 1.4% of step time.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+34
- Incentives
- Insufficient
- Confidence27