Skip to content

Build1 publisher3 min readPublished

The sparse-model bill arrives at serving time, and it is paid in collectives

A CUDA and ROCm tutorial ends on the unglamorous half of Mixture of Experts: two all-to-all collectives per layer, optimizer state pushed onto PCIe, and launch overhead you have to design away.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Part 5 (final part) of the dev.to series "Advanced GPU Optimization: How to tech an LLM with CUDA and ROCm?" states its agenda as: implementing MoE routing and expert parallelism using all-to-all communication, building a ZeRO-Offload mechanism to spill optimizer states to system RAM, harnessing CUDA/HIP graphs to eliminate kernel launch overhead, and building an optimized inference server with KV caching and PagedAttention.
  • The article lists prerequisites as Parts 1-4, a multi-GPU setup (4+ ideal), and "a realization that training is only half the battle - serving is the other half".
  • The author states that in 2024/2025 every major model (naming Grok, Mixtral, Gemini) uses Mixture of Experts to scale to trillions of parameters without exploding compute costs.
  • A dense 1.5 trillion parameter model would require 3TB of VRAM just for weights.
  • MoE solves the VRAM problem by having hundreds of "expert" FFNs but activating only 2 of them per token.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

The final installment of a five-part CUDA and ROCm tutorial series on dev.to spends its length on four mechanisms rather than one architecture: Mixture of Experts routing with all-to-all expert parallelism, a ZeRO-Offload path that spills optimizer state to system RAM, CUDA/HIP graphs to eliminate kernel launch overhead, and an inference server with KV caching and PagedAttention [1]. The author's framing is the useful part: training is only half the battle, serving is the other half, and the prerequisites are Parts 1 through 4 plus a multi-GPU box, four cards or more [2].

Start with the arithmetic that forces the design. A dense 1.5 trillion parameter model would need 3TB of VRAM for weights alone [3], which works out to two bytes per parameter, so that figure already assumes half precision and gets no rescue from further quantisation of the same kind [14]. MoE answers this by holding hundreds of expert feed-forward networks and activating two of them per token [4], with the output a softmax-weighted sum over the top-k experts, k usually 2 [5]. The author notes that in 2024 and 2025 the major models, naming Grok, Mixtral and Gemini, use this to reach trillions of parameters without a matching compute bill [18].

What that buys in FLOPs it spends on the network. Dense models used All-Reduce; MoE needs All-to-All, because different GPUs own different experts and every token has to travel to whichever rank holds the expert it was routed to [8]. The listing is three steps: an NCCL/RCCL grouped send and receive built from per-rank send and receive counts, a local expert kernel over whatever arrived, then a second all-to-all to return the processed tokens to their origin GPUs [9]. That is two collectives per MoE layer per forward pass, and both are shaped by routing decisions that are not known until the router runs [15]. The author asserts this scales to thousands of experts across a cluster with minimal communication overhead [19]; that claim is unquantified in the text and is exactly the number an operator should measure rather than accept.

The router kernel itself is honest about being illustrative. It tracks the top two scores in a manual reduction and applies softmax to only those two, which is the right sparse move [6]. But it runs one thread per token and loops over every expert and every hidden dimension to get the logits [7], so routing cost per token scales with the full expert count regardless of k [c8a].

The offload section targets the other wall. At 70B parameters, ZeRO-3 may not fit in 40GB of VRAM, so momentum, variance and sometimes gradients go to host DDR [10], allocated as pinned memory [12], with gradients copied out asynchronously during the backward pass, the AdamW update computed on a CPU thread, and FP32 weights copied back just in time for the next forward [11]. If those two moment buffers are FP32, as the FP32 master weights imply, they are roughly 560GB on their own, about fourteen times a 40GB card [16]. The throughput ceiling of that scheme is PCIe, not the GPU.

Watch the two items the supplied text promises but breaks off before delivering: the graph capture and the paged KV cache both appear only in the agenda, with the article cutting off mid-listing inside the offload code [17]. Those are the pieces that decide decode latency, and a tutorial that stops before them has described the problem, not the fix.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories