Ai2 has published a detailed account of how it schedules GPU time, and the headline is that it stopped handing out priorities and started handing out budgets. In a post on the Hugging Face blog, the Ai2 infrastructure team describes replacing a priority-based scheduler with GPU time budgets, hierarchical fair-share allocation and a time-slicing contract. Over a 30-day test, teams received 98% of the GPU hours they were owed while cluster occupancy held at 98%.
For anyone running a shared training cluster, this is a rare, numbers-backed look at a problem every lab has: demand for GPUs that far exceeds supply. It also lands in a week when compute is the story everywhere, from AWS raising Capacity Blocks prices to cities pausing new data centers.
TL;DR
| Question | Answer |
|---|---|
| Who | The AI Infrastructure team at Ai2 |
| Scale | Thousands of H100, B200 and B300 GPUs in clusters of 88 to 1,024 GPUs, about 150 researchers |
| Old system | Priority-based scheduler with opt-out preemption |
| New system | GPU time budgets, hierarchical fair-share, minimum-runtime contract |
| Rollout | Cluster by cluster from the end of July 2026 |
| Fairness result | 98% of owed GPU hours delivered; 13 of 15 team allocations at 95% or more; worst case 90% |
| Occupancy | 98% before and after; 18% of delivered time was unallocated |
| Debug jobs | p90 queue wait from 2 hours to 30 seconds |
| Largest H100 cluster | Median wait from 5 minutes to 24 seconds; p90 from 2.8 to 1.8 hours |
| Ops | 74% fewer repairs needing a human in the loop |
Why did the old scheduler break?
Ai2 frames its work as a pyramid of four metrics. Availability is whether hardware is healthy. Occupancy is the fraction of available time assigned to a workload. Impact is how often the most valuable workloads are chosen. Utilization is how much of the GPU a running workload actually uses. The post is about the third layer, impact.
The team says demand sits at 2 to 3 times supply: every available GPU hour has two or three research workloads competing for it. The old design gave each team a limit of concurrent GPUs for non-preemptible workloads, and let preemptible workloads soak up idle GPUs. It produced three predictable pathologies.
- GPU squatting. Researchers could not launch debugging jobs quickly enough, so they parked no-op workloads they could connect to later.
- Priority inflation. Eventually 100% of scheduled workloads used HIGH priority, starving everything lower.
- On-call toil. Because preemptibility was optional, engineers spent most of their ticket time negotiating the shutdown of non-preemptible jobs on hosts with known maintenance problems.
Ai2 admits it was slow to see the root cause. Its first fixes were tighter control over priorities and explicit GPU monopolies for important projects. In hindsight, it says, it had built a laboratory for the tragedy of the commons.
How do GPU time budgets work?
The classic answer to a commons problem is ownership. Giving teams dedicated GPUs, though, left hardware idle because research is seasonal: teams are ready to run at different times. Ai2 compared this to solving a knapsack problem by hand.
Its alternative was to allocate a share of GPU time instead of GPUs. Demand cannot be forecast because it depends on the results of experiments, but priority across research efforts is a strategic call that can be debated in advance. Leadership, in the post's words, can think like investors: fund each effort with GPU time before the workloads exist.
The allocation is hierarchical. A manager divides time among the projects and people under them, so a project like the post's example "A1" knows it holds a fixed claim, 35% in the diagram, regardless of who else is queuing. Allocation decisions happen at the level with the most context: a lead researcher within a project, a principal investigator within a program, a lead program manager or the CEO across programs.
The key rule is that every request for protected GPU time must be funded by a budget. HIGH priority used to be free. Now nothing is, so any trick draws on the beneficiary's own allocation. Ai2's stated strategy is to make gaming the scheduler more expensive than honestly arguing for a bigger budget.
What is the fair-share scheduler doing?
The scheduling algorithm itself is not new, and Ai2 says so. Hierarchical fair-share over a time window traces back to the Hadoop Fair Scheduler in 2009 and lives on in SLURM's Fair Tree and YARN's Fair Scheduler. What is new is the inputs: the tree mirrors the research program structure, and the weights are manager-set budgets rather than static quotas.
The scheduler tracks occupancy over a sliding lookback window, 7 days by default, and sorts workloads from under-used allocations above those from over-used ones. Over a week, any group that keeps submitting enough work should get its share.
It also separates two kinds of occupancy:
| Kind | Charged to a budget | Preemption |
|---|---|---|
| Allocated | Yes, counts toward fair-share | Protected during the minimum runtime |
| Unallocated | No | Always preemptible by any allocated request |
Unallocated time is what keeps occupancy at 98% when funded work is not ready, and it means nobody has a reason to turn down free cycles. In the post's numbers, 18% of delivered GPU time was unallocated. One researcher quoted by Ai2, Chris Clark, said the system feels like an extra 30% of compute, because bursty teams can now reclaim slack and burst past their allocation later without preemption.
What is the scheduling contract?
Training jobs can run for hours, days or weeks, which is the property that made squatting possible. Ai2's fix is a contract: in exchange for cluster access, a workload declares a minimum runtime, the shortest occupancy it needs to make meaningful progress. The workload is protected from preemption for that window. After that, the scheduler may preempt and automatically requeue it if it is resumable. Setting the minimum to zero marks the work as unallocated, free and always preemptible. Ai2 caps the maximum minimum-runtime at 8 hours.
The lifecycle is: submit with a minimum runtime and a resumable flag, get scheduled by fair-share weighted by actual versus allocated occupancy, run through the minimum window (charged to the allocation), keep running while allocations still favor it, possibly be preempted and requeued, then release resources on completion.
An unexpected benefit: unhealthy hosts now drain on their own as workloads reach their minimum runtimes, so repairs can be automated. That cut repairs requiring a human by 74%.
How did Ai2 test the change before rollout?
Because the problem is zero-sum, giving time to one team takes it from another, and losers tend to invent workarounds. Ai2 built a small simulator that takes workloads and submission schedules, then makes preemption and placement decisions, jumping ahead to schedulable moments so it can replay many simulated days in seconds. It used the simulator to tune knobs such as the lookback window and the 8-hour cap.
One hypothesis concerned debug workloads, jobs that need few GPUs and 15 minutes or less. Simulation predicted p90 wait would fall from about 6 hours to 5 minutes. Reality beat the prediction: from 2 hours to 30 seconds. Ai2 notes the baseline had few debug jobs, so that measurement has higher variance.
What happened after rollout?
Rollout began cluster by cluster at the end of July. Over the 30-day period Ai2 reports:
- Teams were delivered 98% of owed GPU hours, with owed time counted as allocation capped hour by hour at actual demand.
- 13 of 15 team allocations received 95% or more, with a worst case of 90%.
- Occupancy stayed at 98% with demand still 2 to 3 times capacity.
- On the largest H100 cluster, median queue wait dropped from 5 minutes to 24 seconds and p90 from 2.8 hours to 1.8 hours.
Against its three original problems: short debug jobs start in under a minute so squatting loses value and costs the squatter budget; priority still exists but only sorts work within a team; and unhealthy hosts drain automatically.
What did not go well?
Ai2 is candid about the costs. The learning curve was steeper than expected, and leftover terminology such as "priority" meant different things after the change. Documentation did not fix the confusion. Live explanatory sessions did, along with new visualizations showing allocation usage and the exact metric used to sort the queue, so a preempted researcher can see why.
The clearest regression is interactive sessions. Researchers used to hold data-analysis and code-testing sessions for up to a week. Under time-slicing they hit the 8-hour protected cap, and preemption meant rebuilding volatile state by hand. After a survey, Ai2 opened two roadmap items: a CPU-only cluster next to on-prem storage for data-prep dev sessions, and restorable sessions so work can resume elsewhere. It is also watching capacity fragmentation, which can lengthen queue waits.
What can other teams take from this?
Ai2's setup is a research institute with about 150 users, but the lessons travel.
- Make priority cost something. Free knobs get maxed out. If HIGH is free, everything becomes HIGH.
- Budget time, not hardware. It keeps the ownership incentive without leaving idle machines.
- Keep a free, preemptible tier. It sustains occupancy without letting anyone avoid spare cycles.
- Require resumability for fairness. Time-slicing only works if jobs can be requeued, which means checkpointing discipline.
- Simulate first, then explain. The simulator was accurate directionally, and live sessions did more for adoption than docs.
- Protect fast debugging. Sub-minute debug starts change developer behavior more than throughput numbers do.
The classic background reading is the Dominant Resource Fairness paper by Ghodsi et al., which Ai2 cites for an anecdote about users sprinkling infinite loops in code to inflate utilization numbers.
Why does this matter beyond Ai2?
Compute scarcity is shaping every part of AI. Labs are chasing gigawatts, cloud providers are repricing capacity, and even frontier labs describe dedicating a share of compute to safety monitoring. When supply is fixed, how it is divided internally is a strategic lever, not an ops detail. A scheduler that gives a program a guaranteed share is effectively an organizational policy encoded in software.
For individual builders squeezing models onto a single card, as in this consumer GPU speed test, the story is different, but the underlying tension is the same: scarce compute rewards whoever wastes the least of it. Power is the other bottleneck, covered in our look at hyperscaler nuclear deals.
Caveats
All figures are self-reported by Ai2 from its own clusters over a 30-day window shortly after rollout, with a small baseline sample for debug jobs. The post does not release the scheduler code, so it is a design description rather than something you can install. Smaller shops may find a full budgeting tree heavier than they need.
Related reading
- AWS Capacity Blocks GPU price hike
- San Francisco data center moratorium
- China 24 GW compute vs US 56 GW
- OpenAI on 5 to 10 percent compute for safety monitoring
- Every hyperscaler nuclear deal
- Qwen3 8B Flash on a consumer GPU
Details reflect Ai2's post of October 9, 2026 and may change.
