Skip to content

Build1 publisher2 min readPublished

Halving vLLM's attention tiles puts Gemma 4 on SageMaker's smallest GPU at 0.8x L4 speed

Gemma 4 decodes on SageMaker's smallest GPU, a T4, at 0.8x an L4's speed with matching outputs from a patched vLLM image, a dev.to benchmark reports. The T4 is the cheaper choice per token only when it rents for under 80% of the L4's hourly rate.

The Engineer · Build desk

Illustration accompanying Halving vLLM's attention tiles puts Gemma 4 on SageMaker's smallest GPU at 0.8x L4 speed

What happened

  • SageMaker JumpStart starts its Gemma 4 listings at ml.g6e.xlarge, a single NVIDIA L40S, and none of its Gemma 4 entries lists a T4.
  • vLLM issue #38918, about Gemma 4 hitting shared memory limits on Turing GPUs, is still open and was reported again on vLLM 0.29.0 in September.
  • The author's patch halves vLLM's attention tiles on pre-Ampere GPUs, and the build log reports a 60,000-byte shared memory budget against Turing's 65,536-byte limit.
  • In the author's 26B A4B build, 0 of 11,534,336 layer-0 query weights leave Google's 4-bit QAT grid, while a June AWQ build of the same export moves 30.3% by a level or more.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Choosing the T4 means maintaining a derived image: every new AWS base container has to be rebuilt with the patch and pass its checks until vLLM fixes Turing upstream.
  • exposure A T4 deployment copied from this series depends on one author's patch script and one author's checkpoint repacks, and the deployer inherits any error in either.
  • capability Because the patch changes nothing on an L4 or newer and keeps the stock entrypoint, a single image can back both T4 and L4 variants with the same SM_VLLM_* settings.

On a T4, stock vLLM stops Gemma 4 at start-up. The model's full-attention layers are 512 wide, and vLLM serves them with its Triton attention kernel, whose tile asks for 98,304 bytes of shared memory per block [5]. Turing allows 65,536 [6]. The kernel wants 1.5 times what the card has [1]. vLLM 0.30.0, the version in the SageMaker container, has no fix [7].

The author's workaround is the tile-halving script already used on a Compute Engine T4 [19], copied into a Dockerfile built FROM the stock container [8]. The build runs the script a second time with --check, and the last RUN line exits with an error if the image's PyTorch has no sm_75 kernels [10]. I like that guard. A base image that drops Turing support fails in CodeBuild, where the failure is cheap to find. CodeBuild also pushes the result to a private ECR repository, so the 8 GB base image never touches the local disk [11], a small mercy for anyone building over hotel Wi-Fi.

The 0.8x ratio is the author's headline figure, set against part three's L4 measurements [1][17]. Both runs served the author's repacks of Google's QAT weights with 4-bit embeddings [13]. For the ratio to carry over to someone else's endpoint, the checkpoints and the patched kernel have to be the same ones. The article's grid check compares the 26B A4B checkpoint with an AWQ build of the same QAT export [14]. The claim that the T4 and L4 give the same answers comes from the author's headline [1].

The write-up does not give hourly prices for either instance. Cost per token is the hourly price divided by tokens per hour. At 0.8x, the T4 makes four tokens for every five the L4 makes, so its hourly price has to sit below 0.8 times the L4's to win [2]. Quota matters too. In the author's account, ml.g4dn.xlarge had an endpoint quota of 2 in each of four US regions [15]. The other small accelerators need a different stack. The Graviton T4G exists on EC2 only, and Inferentia2 needs the Neuron SDK, a different container and a compiled model [16].

One more setting sits below the container. Each SageMaker production variant runs on one of several host images, each with its own NVIDIA driver, chosen by InferenceAmiVersion [18].

What to watch

  • A merged fix for vLLM issue #38918 would let AWS's stock SageMaker container start Gemma 4 on a T4 without the patched image.
  • A T4 (ml.g4dn) entry in JumpStart's Gemma 4 listings would move this setup from self-built to AWS-listed.
  • Published per-build decode rates and current hourly prices for both instances would let the 0.8x ratio be checked as cost per token.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories