explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Why did the old scheduler break?
  • How do GPU time budgets work?
  • What is the fair-share scheduler doing?
  • What is the scheduling contract?
  • How did Ai2 test the change before rollout?
  • What happened after rollout?
  • What did not go well?
  • What can other teams take from this?
  • Why does this matter beyond Ai2?
  • Caveats
  • Related reading
← Back to blog

explainx / blog

Ai2 Rebuilt Its GPU Scheduler Around Time Budgets: 98% of Owed Hours Delivered

AI Infrastructure, Ai2, GPU, Research, Compute

Part of AI Chips and Infrastructure

Ai2 replaced priority queues with GPU time budgets and fair-share scheduling. Debug waits fell from 2 hours to 30 seconds at 98% occupancy.

Oct 9, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
Ai2 Rebuilt Its GPU Scheduler Around Time Budgets: 98% of Owed Hours Delivered

Ai2 has published a detailed account of how it schedules GPU time, and the headline is that it stopped handing out priorities and started handing out budgets. In a post on the Hugging Face blog, the Ai2 infrastructure team describes replacing a priority-based scheduler with GPU time budgets, hierarchical fair-share allocation and a time-slicing contract. Over a 30-day test, teams received 98% of the GPU hours they were owed while cluster occupancy held at 98%.

For anyone running a shared training cluster, this is a rare, numbers-backed look at a problem every lab has: demand for GPUs that far exceeds supply. It also lands in a week when compute is the story everywhere, from AWS raising Capacity Blocks prices to cities pausing new data centers.

TL;DR

table · 2 cols
QuestionAnswer
WhoThe AI Infrastructure team at Ai2
ScaleThousands of H100, B200 and B300 GPUs in clusters of 88 to 1,024 GPUs, about 150 researchers
Old systemPriority-based scheduler with opt-out preemption
New systemGPU time budgets, hierarchical fair-share, minimum-runtime contract
RolloutCluster by cluster from the end of July 2026
Fairness result98% of owed GPU hours delivered; 13 of 15 team allocations at 95% or more; worst case 90%
Occupancy98% before and after; 18% of delivered time was unallocated
Debug jobsp90 queue wait from 2 hours to 30 seconds
Largest H100 clusterMedian wait from 5 minutes to 24 seconds; p90 from 2.8 to 1.8 hours
Ops74% fewer repairs needing a human in the loop
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Why did the old scheduler break?

Ai2 frames its work as a pyramid of four metrics. Availability is whether hardware is healthy. Occupancy is the fraction of available time assigned to a workload. Impact is how often the most valuable workloads are chosen. Utilization is how much of the GPU a running workload actually uses. The post is about the third layer, impact.

The team says demand sits at 2 to 3 times supply: every available GPU hour has two or three research workloads competing for it. The old design gave each team a limit of concurrent GPUs for non-preemptible workloads, and let preemptible workloads soak up idle GPUs. It produced three predictable pathologies.

  • GPU squatting. Researchers could not launch debugging jobs quickly enough, so they parked no-op workloads they could connect to later.
  • Priority inflation. Eventually 100% of scheduled workloads used HIGH priority, starving everything lower.
  • On-call toil. Because preemptibility was optional, engineers spent most of their ticket time negotiating the shutdown of non-preemptible jobs on hosts with known maintenance problems.

Ai2 admits it was slow to see the root cause. Its first fixes were tighter control over priorities and explicit GPU monopolies for important projects. In hindsight, it says, it had built a laboratory for the tragedy of the commons.

How do GPU time budgets work?

The classic answer to a commons problem is ownership. Giving teams dedicated GPUs, though, left hardware idle because research is seasonal: teams are ready to run at different times. Ai2 compared this to solving a knapsack problem by hand.

Its alternative was to allocate a share of GPU time instead of GPUs. Demand cannot be forecast because it depends on the results of experiments, but priority across research efforts is a strategic call that can be debated in advance. Leadership, in the post's words, can think like investors: fund each effort with GPU time before the workloads exist.

The allocation is hierarchical. A manager divides time among the projects and people under them, so a project like the post's example "A1" knows it holds a fixed claim, 35% in the diagram, regardless of who else is queuing. Allocation decisions happen at the level with the most context: a lead researcher within a project, a principal investigator within a program, a lead program manager or the CEO across programs.

The key rule is that every request for protected GPU time must be funded by a budget. HIGH priority used to be free. Now nothing is, so any trick draws on the beneficiary's own allocation. Ai2's stated strategy is to make gaming the scheduler more expensive than honestly arguing for a bigger budget.

What is the fair-share scheduler doing?

The scheduling algorithm itself is not new, and Ai2 says so. Hierarchical fair-share over a time window traces back to the Hadoop Fair Scheduler in 2009 and lives on in SLURM's Fair Tree and YARN's Fair Scheduler. What is new is the inputs: the tree mirrors the research program structure, and the weights are manager-set budgets rather than static quotas.

The scheduler tracks occupancy over a sliding lookback window, 7 days by default, and sorts workloads from under-used allocations above those from over-used ones. Over a week, any group that keeps submitting enough work should get its share.

It also separates two kinds of occupancy:

table · 3 cols
KindCharged to a budgetPreemption
AllocatedYes, counts toward fair-shareProtected during the minimum runtime
UnallocatedNoAlways preemptible by any allocated request

Unallocated time is what keeps occupancy at 98% when funded work is not ready, and it means nobody has a reason to turn down free cycles. In the post's numbers, 18% of delivered GPU time was unallocated. One researcher quoted by Ai2, Chris Clark, said the system feels like an extra 30% of compute, because bursty teams can now reclaim slack and burst past their allocation later without preemption.

What is the scheduling contract?

Training jobs can run for hours, days or weeks, which is the property that made squatting possible. Ai2's fix is a contract: in exchange for cluster access, a workload declares a minimum runtime, the shortest occupancy it needs to make meaningful progress. The workload is protected from preemption for that window. After that, the scheduler may preempt and automatically requeue it if it is resumable. Setting the minimum to zero marks the work as unallocated, free and always preemptible. Ai2 caps the maximum minimum-runtime at 8 hours.

The lifecycle is: submit with a minimum runtime and a resumable flag, get scheduled by fair-share weighted by actual versus allocated occupancy, run through the minimum window (charged to the allocation), keep running while allocations still favor it, possibly be preempted and requeued, then release resources on completion.

An unexpected benefit: unhealthy hosts now drain on their own as workloads reach their minimum runtimes, so repairs can be automated. That cut repairs requiring a human by 74%.

How did Ai2 test the change before rollout?

Because the problem is zero-sum, giving time to one team takes it from another, and losers tend to invent workarounds. Ai2 built a small simulator that takes workloads and submission schedules, then makes preemption and placement decisions, jumping ahead to schedulable moments so it can replay many simulated days in seconds. It used the simulator to tune knobs such as the lookback window and the 8-hour cap.

One hypothesis concerned debug workloads, jobs that need few GPUs and 15 minutes or less. Simulation predicted p90 wait would fall from about 6 hours to 5 minutes. Reality beat the prediction: from 2 hours to 30 seconds. Ai2 notes the baseline had few debug jobs, so that measurement has higher variance.

What happened after rollout?

Rollout began cluster by cluster at the end of July. Over the 30-day period Ai2 reports:

  • Teams were delivered 98% of owed GPU hours, with owed time counted as allocation capped hour by hour at actual demand.
  • 13 of 15 team allocations received 95% or more, with a worst case of 90%.
  • Occupancy stayed at 98% with demand still 2 to 3 times capacity.
  • On the largest H100 cluster, median queue wait dropped from 5 minutes to 24 seconds and p90 from 2.8 hours to 1.8 hours.

Against its three original problems: short debug jobs start in under a minute so squatting loses value and costs the squatter budget; priority still exists but only sorts work within a team; and unhealthy hosts drain automatically.

What did not go well?

Ai2 is candid about the costs. The learning curve was steeper than expected, and leftover terminology such as "priority" meant different things after the change. Documentation did not fix the confusion. Live explanatory sessions did, along with new visualizations showing allocation usage and the exact metric used to sort the queue, so a preempted researcher can see why.

The clearest regression is interactive sessions. Researchers used to hold data-analysis and code-testing sessions for up to a week. Under time-slicing they hit the 8-hour protected cap, and preemption meant rebuilding volatile state by hand. After a survey, Ai2 opened two roadmap items: a CPU-only cluster next to on-prem storage for data-prep dev sessions, and restorable sessions so work can resume elsewhere. It is also watching capacity fragmentation, which can lengthen queue waits.

What can other teams take from this?

Ai2's setup is a research institute with about 150 users, but the lessons travel.

  1. Make priority cost something. Free knobs get maxed out. If HIGH is free, everything becomes HIGH.
  2. Budget time, not hardware. It keeps the ownership incentive without leaving idle machines.
  3. Keep a free, preemptible tier. It sustains occupancy without letting anyone avoid spare cycles.
  4. Require resumability for fairness. Time-slicing only works if jobs can be requeued, which means checkpointing discipline.
  5. Simulate first, then explain. The simulator was accurate directionally, and live sessions did more for adoption than docs.
  6. Protect fast debugging. Sub-minute debug starts change developer behavior more than throughput numbers do.

The classic background reading is the Dominant Resource Fairness paper by Ghodsi et al., which Ai2 cites for an anecdote about users sprinkling infinite loops in code to inflate utilization numbers.

Why does this matter beyond Ai2?

Compute scarcity is shaping every part of AI. Labs are chasing gigawatts, cloud providers are repricing capacity, and even frontier labs describe dedicating a share of compute to safety monitoring. When supply is fixed, how it is divided internally is a strategic lever, not an ops detail. A scheduler that gives a program a guaranteed share is effectively an organizational policy encoded in software.

For individual builders squeezing models onto a single card, as in this consumer GPU speed test, the story is different, but the underlying tension is the same: scarce compute rewards whoever wastes the least of it. Power is the other bottleneck, covered in our look at hyperscaler nuclear deals.

Caveats

All figures are self-reported by Ai2 from its own clusters over a 30-day window shortly after rollout, with a small baseline sample for debug jobs. The post does not release the scheduler code, so it is a design description rather than something you can install. Smaller shops may find a full budgeting tree heavier than they need.

Related reading

  • AWS Capacity Blocks GPU price hike
  • San Francisco data center moratorium
  • China 24 GW compute vs US 56 GW
  • OpenAI on 5 to 10 percent compute for safety monitoring
  • Every hyperscaler nuclear deal
  • Qwen3 8B Flash on a consumer GPU

Details reflect Ai2's post of October 9, 2026 and may change.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 8, 2026

China Has 24 GW of Data Center Capacity vs 56 GW in the US: What SemiAnalysis Found

SemiAnalysis, as relayed by the Financial Times, says China has more than 24 GW of operating data center capacity with roughly 50 GW more planned or announced, against about 56 GW in the US by the end of 2026. Gigawatts are not the same as compute, so here is how to read the figures.

Sep 29, 2026

Nvidia–Anthropic $180B Contracted Value: What Builders Pay

On September 28, 2026, Nvidia said contracted value with Anthropic exceeds $180 billion. That is a Nvidia-specific stack — GPU-backed compute, the November 2025 up-to-$10 billion equity arrangement, and IPO-anchor talks — not Anthropic's reported ~$517 billion multi-vendor compute total. This post maps the stack to Claude API capacity, rate limits, and vendor risk.

Sep 10, 2026

OpenAI Reportedly Plans $750 Billion in Compute Spend Through 2030

OpenAI reportedly plans to spend roughly $750 billion on compute infrastructure through 2030 — and is still reported to be short on capacity relative to demand. explainx.ai breaks down what this figure actually represents, how it compares to Anthropic's own reported $517 billion compute commitment, and why this connects directly to the usage limits GPT-6 Astra users have already been hitting.