Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Ai2 swaps priority scheduling for GPU time budgets after every job ended up marked HIGH

Ai2 replaced priority scheduling on thousands of GPUs with time budgets and hierarchical fair-share after every scheduled workload ended up marked HIGH. The change moves the argument over who gets GPU time out of case-by-case operations and into an administrative budget.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Ai2 swaps priority scheduling for GPU time budgets after every job ended up marked HIGH
Generated illustration

What happened

  • At any moment Ai2 has outstanding requests for two to three times more GPUs than it has available, so each GPU hour has two or three workloads competing for it.
  • On-call engineers spent most of their ticket response time negotiating shutdowns of non-preemptable jobs on hosts with known maintenance problems.
  • Ai2 first tried tighter control over how priorities were set, then worked around its own scheduler by handing GPU monopolies to important projects.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A lab copying this design has to decide first whether workloads may opt out of preemption, because at Ai2 that option is what turned host maintenance into negotiation.
  • cost At 2-3x oversubscription, half to two-thirds of requested capacity waits under any scheduler, so a budget only changes which teams carry the wait.
  • precedent Disputes over allocation move to whoever sets the budgets, so the people who own that process become responsible for research priorities.

A priority field orders work only while some jobs sit below the top level. At Ai2, priority inflation ran until 100% of scheduled workloads used HIGH, and the lower levels got no GPU time at all [8]. At that point the priority field ranked nothing. The preemption rule added its own incentive. A team's concurrency cap covered only jobs protected from preemption, and preemptible jobs could exceed it on idle GPUs [6].

The first fixes stayed inside the old model. The team tightened control over how priorities were set, then worked around its own scheduler by assigning GPU monopolies to important projects [10]. Its engineers now call the period a "tragedy of the commons", in which researchers maximizing their own results on a shared resource produced a worse global result [13]. Monopolies are the classic answer to a commons: privatize it [13]. The available text of the post ends mid-sentence in that passage [14]. It does not describe how budgets are sized or what the time-slicing contract requires, and it includes no before-and-after figures.

Gaming of this kind is old. In the 2011 paper that introduced Dominant Resource Fairness, Ghodsi et al. recount a search company that gave dedicated machines only to jobs whose users could guarantee high utilization. The company, they wrote, soon discovered "users would sprinkle their code with infinite loops to artificially inflate utilization levels" [12]. Ai2's version was a no-op workload parked on a GPU for later [7].

Ai2 ranks its infrastructure goals as a pyramid of availability, occupancy, impact and utilization [11]. Occupancy is the share of available time assigned to a workload [11]. A parked no-op job counts toward it [7]. The budget redesign targets impact, defined as how often the most valuable workloads are chosen to receive resources [11].

For another lab, the question is which of these failures it shares. Ai2 has outstanding requests for 2-3x the GPUs it has [5]. At that ratio, between half and two-thirds of requested capacity waits at any moment under any scheduler [15]. Budgets decide whose work waits. They do not address the cause Ai2 gave for squatting, which was that debugging workloads could not launch with low enough latency [7]. I'd expect a lab that copies the budgets without a fast interactive launch path to see parked jobs come back.

The on-call cost is the strongest case for copying. Engineers spent a majority of their ticket response time negotiating the shutdown of non-preemptable workloads on hosts with known maintenance problems [9]. That time traces directly to letting workloads opt out of preemption [6]. A lab weighing this design has to settle that rule before it sets a single budget.

Scale and governance matter too. About 150 researchers share thousands of H100, B200 and B300 GPUs in clusters of 88 to 1,024 [3][4]. Ai2 describes the new allocation as a transparent administrative budgeting process [2]. Someone has to hold the authority to run it.

What to watch

  • Whether Ai2 publishes before-and-after figures for impact, squatting or on-call ticket time under the budget system.
  • The terms of the time-slicing contract: whether preemption becomes mandatory and how much notice a workload gets before it yields a GPU.
  • Whether Ai2 adds a low-latency launch path for debugging jobs, the cause it gave for GPU squatting.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence35
Adoption20
Hype gap+10
Incentives35
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Ai2's AI Infrastructure team replaced a priority-based GPU scheduler with a system including GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract.

    ReportedSupportedSource: Ai2 AI Infrastructure team, Hugging Face blogView cited source
  2. [2]

    Ai2 says the change shifted the debate about how much GPU time each research project deserves from a case-by-case operational task to a transparent administrative budgeting process.

    ReportedSupportedSource: Ai2 AI Infrastructure teamView cited source
  3. [3]

    Ai2 manages thousands of NVIDIA H100, B200 and B300 GPUs in clusters ranging from 88 to 1024 GPUs, built for large-scale distributed training.

    ReportedSupportedSource: Ai2 AI Infrastructure teamView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. huggingface.co

    1 article · October 9, 2026

    Impactful scheduling for GPU clusters

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories