buildOne report1 publisher SageMaker costs 1.40x a plain EC2 instance per hour to serve one Gemma 4 vLLM build on the same T4 or L4 GPU, a benchmark on dev.to finds. With decode speed matched within 2%, the extra 40% goes to the managed layer around the GPU.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence60
buildOne report1 publisher Gemma 4 decodes on SageMaker's smallest GPU, a T4, at 0.8x an L4's speed with matching outputs from a patched vLLM image, a dev.to benchmark reports. The T4 is the cheaper choice per token only when it rents for under 80% of the L4's hourly rate.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence40
buildOne report1 publisher A dated pricing run puts the Graviton2-hosted G5g 20 percent below the Intel G4dn per hour, but that host resolves a container image with no kernels for its own T4G, so it compiles vLLM before serving a token.
Reality
- Evidence58
- Adoption27
- Hype gap+16
- Incentives44
- Confidence54
buildOne report1 publisher NVIDIA documents MIG as compute-only, so headless Chromium inside an H100 slice opens no GPU context at all and quietly falls back to CPU rasterising, which surfaces as a frame rate problem rather than a missing driver.
Reality
- Evidence62
- Adoption26
- Hype gap−8
- Incentives18
- Confidence55
buildOne report1 publisher A 7B model split across Iowa and Oregon on free T4s went from 4.92 to 28.10 tokens per second. Most of the gain came from a drafter that stopped launching kernels one at a time.
Reality
- Evidence42
- Adoption16
- Hype gap+22
- Incentives58
- Confidence38