Ai2 replaces priority scheduling with GPU time budgets: occupancy stays at 98% while debug-job p90 queue wait drops from 2 hours to 30 seconds
A write-up from real clusters on making it costly to game the scheduler.
// Key points
- Ai2's clusters range from 88 to 1024 GPUs and serve about 150 researchers, with demand running 2-3x supply.
- Managers now set GPU time budgets enforced by a hierarchical fair-share scheduler: a declared minimum runtime is protected from preemption and charged to the budget, while unbudgeted time is free but preemptible.
- Occupancy held at 98% before and after; over a 30-day test, teams received 98% of their owed GPU hours, with 13 of 15 allocations at 95% or more.
- Debug-job p90 queue wait fell from 2 hours to 30 seconds, and human-in-the-loop repairs dropped by 74%.
Builder's takeThere's no released code, but the idea transfers directly: swap priority for budgets and grabbing resources gets a cost. When a team shares GPUs for inference or fine-tuning, I'd set budgets per project and keep a fast lane for debug jobs, which beats endlessly tuning priorities.