Build1 publisher3 min readPublished
SageMaker charges 1.40x EC2's hourly rate for the same Gemma 4 server on the same GPU
SageMaker costs 1.40x a plain EC2 instance per hour to serve one Gemma 4 vLLM build on the same T4 or L4 GPU, a benchmark on dev.to finds. With decode speed matched within 2%, the extra 40% goes to the managed layer around the GPU.
The Engineer · Build desk

What happened
- On the T4, both sides held 2.86 GiB of weights and about 660,000 tokens of KV cache, and all 40 test answers came back byte-identical.
- The L4 run, on Google's QAT checkpoint and the stock vLLM 0.30.0 image, showed the model server running at the same speed on both paths.
- The author's timing fit put a fixed cost of 0.56 seconds on each SageMaker call against under 0.01 seconds on the EC2 instance.
- That per-call cost cut SageMaker's throughput at 1 to 16 parallel requests by 1.23x to 1.38x across the two GPUs.
- The EC2 instance was serving 9.7 minutes after launch, with no inbound security-group rules and Systems Manager as its only way in.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Sizing a migration off the 1.23x to 1.38x throughput gap counts client overhead the on-box EC2 test never paid, so the saving on real traffic is likely smaller by an unmeasured amount.
- cost Moving to EC2 to save the 40% hands the team a patched image build, an access path and instance lifecycle code that someone must now maintain.
- decision Because the premium is the same 1.40x on the T4 and the L4, the endpoint-or-instance choice can be made separately from the choice of GPU.
That 0.56 seconds bundles four costs: starting the aws CLI, signing the request, the network round trip from the author's workstation, and the SageMaker front end [8]. The EC2 client ran on the instance beside vLLM, so it paid none of the first three [8]. The author notes that a client off the instance would pay its own network cost and a long-lived SDK client would skip CLI start-up. Neither case was measured, so the front end's own share of the 0.56 seconds is unknown [9].
Taken at face value, the per-call cost compounds the hourly premium. At 16 parallel requests on the T4, EC2 served 1,005.1 tokens a second and SageMaker 787.05 [5]. That 1.28x throughput ratio times the 1.40x hourly ratio gives EC2 about 1.79x the tokens per dollar [2]. The figure carries over to another team only under three conditions. Its callers carry the same overhead as a CLI started on a workstation. Its responses are short enough for a fixed cost to show. It pays the on-demand list prices the comparison used [2]. Under the author's own model, a longer response spreads the same fixed cost over more tokens, decoded at a rate that already matches [7].
The hourly ratio is the clean number. It came from the AWS Price List API on 2026-09-30 and held at 1.40x on both GPUs [2]. The author also priced Google Cloud options from its billing catalog: a T4 on an n1-standard-2 in us-west2, an L4 on a g2-standard-4 in us-east4, and a Cloud Run L4 with 8 vCPU and 32 GiB in us-east4 [16].
The EC2 side is careful work. The model, context length, memory setting, data type and measurement script were held the same on both sides [3]. The rig's settings pin it: `DTYPE=float16`, `MAX_MODEL_LEN=8192`, `GPU_MEMORY_UTILIZATION=0.90` on a `g4dn.xlarge` [11]. Cloud-init pulls the stock vLLM 0.30.0 image, applies the same Turing patch as the series' SageMaker T4 build and builds the derived tag on the instance. The pull finished at +170 seconds and serving started at +211 [12]. The measurement script arrives over Systems Manager and runs against localhost:8000 [15]. Few benchmark rigs plan for their own driver dying; this one has a watchdog that terminates the instance when it does, on top of a `finally` block that terminates it on every normal exit [15].
Those pieces are what a team writes and maintains once it drops the endpoint: a patched image build, an access path and instance lifecycle code [15]. In my view the extra 40% [1] is justified only by operational work of that kind, because on this evidence the GPU does identical work on both paths [4]. The benchmark measured decode speed, 1 to 16 parallel requests and 40 questions at temperature 0 [17]. It did not put a figure on the operator's time.
What to watch
- A rerun with a long-lived SDK client calling the SageMaker endpoint, which would isolate the front end's share of the 0.56 seconds.
- An EC2 run with the client off the instance, adding the network cost the on-box client skipped.
- Pricing under committed-use or savings-plan rates, where the on-demand 1.40x ratio may not hold.