build1 publisher
A tile-size clamp in vLLM's Triton kernel gets Gemma 4 running on a Tesla T4
Getting Gemma 4 E2B to serve on a 2019-era Tesla T4 under vLLM 0.29.0 turns on a clamp inside the Triton attention kernel, and the QAT checkpoint's advantage shows up in host memory before it shows up in tokens per second.
Publishers:dev.to
Reality
- Evidence64
- Adoption14
- Hype gap+22
- Incentives30
- Confidence56