Skip to content

Build1 publisher2 min readPublished

Helion's autotuned GEMM outruns vLLM's CUTLASS and DeepGEMM paths on Hopper GPUs

Helion's autotuned GEMM beat vLLM's default CUTLASS and DeepGEMM backends on Hopper GPUs, by more than 10% throughput on some workloads, its authors report. The gain rests on per-shape tuning that can run for hours, a cost teams pay in place of kernel maintenance.

The Engineer · Build desk

Illustration accompanying Helion's autotuned GEMM outruns vLLM's CUTLASS and DeepGEMM paths on Hopper GPUs
Generated illustration

What happened

  • A single Helion GEMM implementation covers Standard GEMM, Split-K and Swap-AB, and per-shape autotuning picks the variant and its config.
  • vLLM's quantized linear layers in FP8, INT8, INT4 and NVFP4 currently run on backends that wrap kernels from CUTLASS, DeepGEMM and FlashInfer.
  • CUDA Graph capture at vLLM startup triggers Helion JIT compilation and slows cold starts, though cached compiled artifacts largely remove the delay on warm starts.
  • The authors also list a runtime cost from Helion kernel dispatch and launch when inference runs outside the CUDA Graph capture range.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision For a quantized GEMM path, the choice is now between paying for kernel expertise and scheduling tuning runs that can take hours for each deployment.
  • cost Teams that adopt Helion have to ship a compiled-kernel cache with each deployment, or pay for JIT compilation on every cold start.
  • constraint The throughput evidence covers Hopper only. Teams on other GPUs would be adopting on the strength of the portability claim alone.

The phrase that decides how far the result reaches is "hybrid dispatch." According to the PyTorch post, the Hopper win comes from per-shape tuning combined with hybrid dispatch [1]. The post's overview does not define the term. Suppose it means Helion serves the shapes where it wins and the existing backends keep the rest. Then the measured system is Helion plus the incumbents, tested against the incumbents alone. Under that reading, the result supports adding Helion to vLLM's dispatch table, not deleting CUTLASS from it.

The engineering under the result is good. Helion's tuner runs ahead of time and searches one space. That space runs from memory layout and kernel scheduling up to the choice of algorithm [5]. The authors frame tuning as a structured numerical optimization problem. They say the search does not rely on profiling to guide it, and they set it against agentic loops that generate, profile and refine kernels [10]. LLM-guided search is an optional extra mode [11]. Because the algorithm variant sits inside the search space, a new shape gets a new config and not a new kernel [2].

Read the throughput figure as a claim about the evaluated models on Hopper [1]. The more-than-10% gain applies to "some workloads" [1]. To see a comparable gain, you need Hopper hardware and you need to tune on your own shapes. The authors argue the second point themselves. In their account, vLLM's default kernels are tuned for general workloads and common models, and Helion is meant to let users tune for a specific deployment without kernel expertise [9].

The tuning bill is counted in hours. The authors wrote that "fine-grained tuning can still take hours" [13]. They blame tuning granularity, not Helion, and say other DSLs would cost as much or more to reach the same level of specialization [6]. The argument is fair, and it is also one that a tool's authors are well placed to make. The Helion team says work to reduce the integration costs is ongoing [12].

I think the trade favors Helion for a team serving a small, stable set of models on Hopper. Tuning happens ahead of time [5], so the hours go into a build job and stay off the serving path. A team that changes models every week would rerun that job every week. For teams on GPUs other than Hopper, this post has no measurement to offer [1].

What to watch

  • A definition of hybrid dispatch that shows which shapes Helion serves and which stay on CUTLASS or DeepGEMM.
  • Throughput results on GPUs other than Hopper, the first real test of Helion's portability claim.
  • Whether the Helion team's mitigation work reduces dispatch and launch overhead outside the CUDA Graph capture range.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories