Build1 publisher3 min readPublished
Modal's gang scheduler places multi-node GPU training jobs from a shared serverless pool
Modal made multi-node GPU clusters generally available on October 1st, requested through one Python decorator and billed by the second. Dropping a reserved cluster for it means trusting a gang scheduler to place every node at once on one network.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Modal's example, @modal.clustered(size=4, rdma=True), requests four nodes with eight B300 GPUs each, and the API hands every container its rank and the cluster's private IP addresses.
- A new gang scheduler treats each cluster as a single scheduling unit, evaluating pending jobs across Modal's whole fleet before telling the available nodes to start.
- Modal's gVisor runtime lacked the RDMA operations it needed, so the company built a proxy for RDMA verbs and GPU memory access and got it merged upstream.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams that keep reserved GPU clusters running between jobs now have a per-second option to price, and the comparison turns on start times at the cluster size they actually use.
- constraint A team's start time on Clusters depends on how busy Modal's other customers keep the same zone and network, a variable the owner of a reserved cluster does not face.
- capability With the RDMA proxy merged into gVisor itself, other operators running gVisor sandboxes can start from upstream code when they want to give tenants RDMA.
A conventional scheduler can place jobs one machine at a time. Distributed training cannot begin until every node it asked for is placed together [15]. Modal's four-node example therefore needs 32 B300s free at the same moment [1]. According to Modal, the scheduler also keeps the gang inside one availability zone and network [5]. That shrinks the set of free nodes that qualify for any one request.
Clusters draw on the same capacity as Modal's other workloads, according to the company [6]. I think sharing one pool is the right design for utilisation. It is what lets Modal bill by the second with no cluster held between jobs [7]. When the pool runs short, Modal says the scheduler adds capacity [5]. Co-founder and CEO Erik Bernhardsson said in a 2022 account of starting Modal that he wanted software that could take code on a developer's computer and launch it in the cloud within a second [14]. For a gang job, launch time also includes finding every node.
RuntimeWire describes the bet as making distributed workloads feel like ordinary cloud functions without removing the operational demands of keeping a cluster reliable [18]. Modal did not publish start times by cluster size, or say what happens to a running job when one node in the gang fails.
The two GLM 4.7 timings [9] check out as division. A 717 GB payload is 5,736 gigabits. At 50 Gbps that takes about 115 seconds, the "nearly two minutes" in the example [2]. Finishing in under two seconds needs at least 2.9 Tbps of effective throughput, about 45% of the 6.4 Tbps per node that Modal advertises [3]. RDMA matters when GPUs repeatedly synchronize weights during training or transfer large caches during inference [17]. For a gap that size to show up in a team's own job, its wall-clock time has to be dominated by node-to-node transfers of that kind.
I'd rate the gVisor work as the best engineering in the release. Direct access to networking hardware improves performance, while a provider sharing machines among customers still needs isolation between them [16]. Modal extended its sandbox to carry RDMA traffic [10], so the fast path runs inside the same secure, multi-tenant containers as the rest of the platform [19]. The scheduler is only one piece of the stack: the RDMA setup is configured for PyTorch and NCCL, and Clusters work with Modal's Volumes, Cloud Bucket Mounts and Queues [4].
Among its customer examples, Modal says Decagon fine-tuned open models of up to one trillion parameters with the Miles framework [13]. The company raised $355 million in May at a $4.65 billion post-money valuation, in a round led by General Catalyst and Redpoint Ventures [11]. It reported annualized revenue above $300 million [12].
What to watch
- Modal publishing start-time figures for large gang requests, and how it restarts a job when one node of a running cluster fails.
- An independent measurement of a full NCCL training step on Clusters, set against Modal's GLM 4.7 transfer illustration.
- Published per-second prices for B300 clusters that teams can compare with hourly reservation rates.