BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Ai2 swaps priority scheduling for GPU time budgets after every job ended up marked HIGH
Ai2 replaced priority scheduling on thousands of GPUs with time budgets and hierarchical fair-share after every scheduled workload ended up marked HIGH. The change moves the argument over who gets GPU time out of case-by-case operations and into an administrative budget.
The Engineer · Build desk

What happened
- At any moment Ai2 has outstanding requests for two to three times more GPUs than it has available, so each GPU hour has two or three workloads competing for it.
- On-call engineers spent most of their ticket response time negotiating shutdowns of non-preemptable jobs on hosts with known maintenance problems.
- Ai2 first tried tighter control over how priorities were set, then worked around its own scheduler by handing GPU monopolies to important projects.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A lab copying this design has to decide first whether workloads may opt out of preemption, because at Ai2 that option is what turned host maintenance into negotiation.
- cost At 2-3x oversubscription, half to two-thirds of requested capacity waits under any scheduler, so a budget only changes which teams carry the wait.
- precedent Disputes over allocation move to whoever sets the budgets, so the people who own that process become responsible for research priorities.
A priority field orders work only while some jobs sit below the top level. At Ai2, priority inflation ran until 100% of scheduled workloads used HIGH, and the lower levels got no GPU time at all [8]. At that point the priority field ranked nothing. The preemption rule added its own incentive. A team's concurrency cap covered only jobs protected from preemption, and preemptible jobs could exceed it on idle GPUs [6].
The first fixes stayed inside the old model. The team tightened control over how priorities were set, then worked around its own scheduler by assigning GPU monopolies to important projects [10]. Its engineers now call the period a "tragedy of the commons", in which researchers maximizing their own results on a shared resource produced a worse global result [13]. Monopolies are the classic answer to a commons: privatize it [13]. The available text of the post ends mid-sentence in that passage [14]. It does not describe how budgets are sized or what the time-slicing contract requires, and it includes no before-and-after figures.
Gaming of this kind is old. In the 2011 paper that introduced Dominant Resource Fairness, Ghodsi et al. recount a search company that gave dedicated machines only to jobs whose users could guarantee high utilization. The company, they wrote, soon discovered "users would sprinkle their code with infinite loops to artificially inflate utilization levels" [12]. Ai2's version was a no-op workload parked on a GPU for later [7].
Ai2 ranks its infrastructure goals as a pyramid of availability, occupancy, impact and utilization [11]. Occupancy is the share of available time assigned to a workload [11]. A parked no-op job counts toward it [7]. The budget redesign targets impact, defined as how often the most valuable workloads are chosen to receive resources [11].
For another lab, the question is which of these failures it shares. Ai2 has outstanding requests for 2-3x the GPUs it has [5]. At that ratio, between half and two-thirds of requested capacity waits at any moment under any scheduler [15]. Budgets decide whose work waits. They do not address the cause Ai2 gave for squatting, which was that debugging workloads could not launch with low enough latency [7]. I'd expect a lab that copies the budgets without a fast interactive launch path to see parked jobs come back.
The on-call cost is the strongest case for copying. Engineers spent a majority of their ticket response time negotiating the shutdown of non-preemptable workloads on hosts with known maintenance problems [9]. That time traces directly to letting workloads opt out of preemption [6]. A lab weighing this design has to settle that rule before it sets a single budget.
Scale and governance matter too. About 150 researchers share thousands of H100, B200 and B300 GPUs in clusters of 88 to 1,024 [3][4]. Ai2 describes the new allocation as a transparent administrative budgeting process [2]. Someone has to hold the authority to run it.
What to watch
- Whether Ai2 publishes before-and-after figures for impact, squatting or on-call ticket time under the budget system.
- The terms of the time-slicing contract: whether preemption becomes mandatory and how much notice a workload gets before it yields a GPU.
- Whether Ai2 adds a low-latency launch path for debugging jobs, the cause it gave for GPU squatting.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption20
- Hype gap+10
- Incentives35
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Ai2's AI Infrastructure team replaced a priority-based GPU scheduler with a system including GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract.
- [2]
Ai2 says the change shifted the debate about how much GPU time each research project deserves from a case-by-case operational task to a transparent administrative budgeting process.
- [3]
Ai2 manages thousands of NVIDIA H100, B200 and B300 GPUs in clusters ranging from 88 to 1024 GPUs, built for large-scale distributed training.
- [4]
The clusters serve about 150 internal researchers working on LLM and VLM training, robotics RL simulation, and post-training for scientific agentic use cases.
- [5]
Based on submitted workloads, at any moment Ai2 has outstanding requests for 2-3x more GPUs than are available; every available GPU hour has 2-3 research workloads competing for it.
- [6]
Ai2's old priority-based scheduler let workloads opt out of preemptability; each team had a limit on concurrent GPUs for workloads protected from preemption, and preemptible workloads could exceed that limit on idle GPUs.
- [7]
Ai2 observed GPU 'squatting', where users parked no-op workloads they could connect to when needed, because researchers could not launch debugging workloads with low enough latency to tackle problems in real time.
- [8]
Ai2 observed priority inflation where eventually 100% of scheduled workloads used HIGH priority, so lower priority levels were starved of GPU time altogether.
- [9]
Because preemptability was optional, Ai2's on-call engineers spent a majority of their ticket response time negotiating the organized shutdown of non-preemptable workloads running on hosts with known maintenance problems.
- [10]
Ai2's initial attempts focused on tighter control of how priorities were set and, ultimately, working around the priority-based scheduler by explicitly assigning GPU monopolies to important projects.
- [11]
Ai2 frames its work as a pyramid of four metrics: availability (hardware healthy and ready), occupancy (fraction of available time assigned to a workload), impact (how often the most valuable workloads are chosen to receive resources), and utilization (fraction of GPU capacity used over a workload's lifetime). The post is about improving impact.
- [12]
users would sprinkle their code with infinite loops to artificially inflate utilization levels
ReportedSupportedSource: Ghodsi et al., 2011 Dominant Resource Fairness paper, recounting a search company that gave dedicated machines only to jobs whose users could guarantee high utilization; quoted in the Ai2 postView cited source - [13]
Ai2 describes its situation as a 'tragedy of the commons', with individuals competing over a scarce shared resource and achieving a non-optimal global result; it says the classic solution is to privatize the shared resource.
- [14]
The supplied text of the post ends mid-sentence in the passage on assigning teams monopolies over sets of GPUs.
- [15]
At 2-3x oversubscription, between 50% and about 67% of requested GPU capacity is waiting at any moment.
Sources
1 independent publisher whose own reporting we read for this story.
- huggingface.coImpactful scheduling for GPU clusters
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.