Skip to content

Build1 publisher3 min readPublished

Repacked 4-bit embeddings lift Gemma 4 decode up to 1.39x on a single L4

Repacking Gemma 4's QAT weights with 4-bit embedding tables made decode up to 1.39x faster on one SageMaker L4, according to a dev.to benchmark series. For teams serving Gemma 4 on vLLM, how the weights are stored becomes a setting to measure alongside model size.

The Engineer · Build desk

Illustration accompanying Repacked 4-bit embeddings lift Gemma 4 decode up to 1.39x on a single L4

What happened

  • Google's -qat-q4_0-unquantized files store Gemma 4's QAT weights at 16-bit, though every group of 32 values is already a scale times an integer from -8 to 7.
  • Both model configs set tie_word_embeddings to true, yet Google's -w4a16-ct files still ship a separate lm_head.weight, and vLLM loads that copy.
  • A layer the repack leaves at 16-bit that the compressed-tensors ignore list does not name makes vLLM look for packed 4-bit tensors that are not there, so loading fails.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Operators who want 4-bit embeddings, or a W4A16 build of the 26B, on vLLM have to run and verify the repack themselves, because Google's exports provide neither.
  • capability E2B and E4B deployments can hand up to 1.25 GiB of L4 memory back to serving by changing only the files, since decode speed did not move.
  • cost Skipping a pre-deploy diff of the checkpoint index turns each ignore-list slip into a roughly 30-minute wait for SageMaker to mark the endpoint Failed.

The repack tool checks its own work. It recovers each group's integers and writes them as compressed-tensors W4A16, the format vLLM reads, and it verifies every group as it goes [5]. On the per-layer embedding table it turned 4.375 GiB of bf16 into 1.230 GiB of int4 plus f16 scales, with zero off-grid groups [6]. The table ends up at 28% of its original size [1].

Some values do change. Only 74.26% came back bit-identical, and the worst error was 6.58e-03 of the group maximum [6]. Every integer landed on the grid, so I'd expect the residual to come from the scales, which are stored at f16 where the source was bf16. Any damage from it would show in the measurement script's 40 exact-answer questions, asked at temperature 0 [18].

The duplicate lm_head costs memory and buys nothing. The config says the two tables are one; vLLM believes the files instead [9]. The repack leaves the copy out. Decode speed held at 107.2 against 105.1 tok/s on E2B and 60.4 against 60.7 on E4B [10]. The gaps are 2.0% and 0.5% [2].

The speed gain needs one more step. vLLM ties the output layer by copying the embedding's 16-bit weight, and a packed table has no such weight, so the emb4 build stores lm_head as its own 4-bit table [13]. The same repack also yields a W4A16 build of the 26B, a size Google's -qat-w4a16-ct line does not cover [3][7]. The published excerpt stops before the per-build results, so the 1.39x cannot be matched here to a size or a format [1]. I would look at E2B first, the size whose Google file is mostly embedding tables [8].

Google's own configs show what a correct ignore list looks like: 17 entries for the 12B and 250 for E2B and E4B, with 140 of those from the audio tower alone [16]. The failure is at least specific. The error names the 16-bit vision_embedder.patch_dense.weight tensor and lists the packed parameters vLLM built for that layer in its place [15]. The author's pre-deploy check diffs the checkpoint index against the ignore list [17]. I think that script is the most reusable piece of the three-part series [21].

Every build ran on its own ml.g6.xlarge in us-east-2 with max_model_len 8192 and 90% GPU memory, one after the other [20]. All are text-only, with the vision and audio towers removed [11]. The decode rate cancels the aws CLI start-up and network time [19]. The 1-, 4- and 16-request throughput figures include that cost, and it ranged from 0.532 s to 1.232 s per call across runs [19], a 2.3x spread [3].

So the 1.39x is a decode-rate result on one L4 with network time removed. It should transfer to a text-only L4 deployment serving replies near the 16- and 512-token lengths the script uses [18]. Image or audio input, or contexts past 8192 tokens, need their own run. The 8-bit rows also rely on the L4 running both 8-bit formats natively on its tensor cores [12]. On a fixed instance, faster decode means more output per instance-hour. The builds are on Hugging Face under xbill9/ [14], so a rerun costs one endpoint per format.

What to watch

  • The full per-build results table: which Gemma 4 size and storage format produced the 1.39x, and how each build scored on the 40 exact-answer questions.
  • Whether Google adds 4-bit embedding tables or a 26B W4A16 export to its -qat-w4a16-ct line, or drops the duplicate lm_head.weight from those files.
  • Parallel throughput at 4 and 16 concurrent requests measured without the aws CLI's per-call overhead, which would show whether the decode gain survives batching.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories