Build1 publisher3 min readPublished
Repacked Gemma 4 QAT weights run a 12B model at bf16 accuracy on one TPU v5e chip
Repacking Google's quantization-aware-trained Gemma 4 weights without re-rounding fits the 12B model on one TPU v5e chip at bf16-level accuracy. Adopting it means carrying patches to vLLM's TPU backend and working inside a 9,728-token KV cache.
The Engineer · Build desk

What happened
- The same repack method serves every Gemma 4 size from E2B to 26B on a single v5litepod-1 chip.
- The repacked builds score up to 2.4 points above Google's own 4-bit exports on the author's classification suite, at the same speed.
- At 16 parallel requests the 12B int8 build serves 675 output tokens per second and scores 0.964 on GSM8K and 0.955 on BFCL tool calling.
- Every per-record output, log and script is committed, and the builds are published on Hugging Face under xbill9/gemma-4-*-it-qat-*.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- contradiction The write-up names 12B the chip's largest model even though 26B loads, so 'largest' holds only as largest with cache room for one full 4,096-token request.
- decision Picking int4 over int8 for 12B on v5e costs throughput: the 4-bit builds deliver about 58% of the int8 build's output tokens per second.
- cost Anyone adopting the setup owns its patch set, so each vLLM upgrade means re-applying the tpu_inference changes and re-running the accuracy suites.
Quantization-aware training leaves Gemma 4 with weights that already sit on a 4-bit grid, 16 levels for every group of 32 [3]. Google publishes those trained values as bf16 checkpoints with "unquantized" in the name [3]. The name describes the file format and little else. Google's own -qat-w4a16-ct exports re-round every group [17]. Keeping the trained values is the correct export, and the repack does it: it recovers each group's trained step and stores the group as int4, so every value keeps its place on the grid [4]. The int8 builds take the same trained values to int8 per channel [5].
One v5litepod-1 chip has 15.75 GiB of HBM [1]. E4B at bf16 needs 14.9 GiB of that, leaving 0.85 GiB [2][1]. The 12B at bf16 needs 22.4 GiB [2]. The 12B build that fits pairs int8 weights with int4 embedding tables. A patch keeps those tables packed on the chip and unpacks only the rows a step reads [7]. Its weights take 11.31 GiB, so 4.44 GiB remains for everything else [22][2]. vLLM turns that into a 9,728-token KV cache, or 2.38 concurrent requests at 4,096 tokens each [15].
The write-up says the repacks "make 12B the largest model on the chip" [13]. The 26B mixture-of-experts repack loads at 13.58 GiB, leaving 2.17 GiB, and gets a 2,176-token cache [16][4]. That holds about 53% of one 4,096-token request [3]. In my view the 26B result proves the weights load, and 12B is the size to serve. The 26B build is also the one size that reads below bf16, by 1.1 points [10].
On this chip, speed follows the integer type. The v5e multiplies int8 by int8 natively [6]. Between the two 4-bit formats, throughput is a tie. Served back to back on one VM at 16 requests, Google's E2B export ran 1,910.8 output tokens per second and the repack 1,910.2 [18]. The E4B pair ran 1,020 and 1,009, and the 4-bit 12B pair ran 392 and 388 [18].
The accuracy figures describe three workloads. Classification is 3,880 public records from Bespoke Labs' benchmark set, paired record for record against bf16 with 95% ranges [11]. Math and tool calling are GSM8K and BFCL's simple split, scored in code by the author's gen_eval.py against the served model [20]. For E4B, 12B and 26B, the bf16 references ran on v6e, because those sizes do not fit the v5e at bf16 [19]. The parity result transfers to traffic that resembles those tasks and fits inside 9,728 tokens of cache [15].
Two more additions to vLLM's tpu_inference complete the serving path: an int8 W8A8 method on the JAX path and an int4 lm_head [7]. The author's runs apply all three patches at boot to a pinned vLLM image [8]. On that setup the 12B int8 build reported ready 405 seconds after it started loading [23].
What to watch
- Whether the int8 W8A8, packed int4 embedding and int4 lm_head changes appear in a released vLLM TPU image, removing the pinned-image step.
- Whether Google changes its -qat-w4a16-ct exports to keep the trained grid values, which would close the gap of up to 2.4 points.
- Long-prompt results for the 26B build, whose cache holds 2,176 tokens.