Research Oct 10, 2026 · Originally published Oct 9

Ai2 replaces priority scheduling with GPU time budgets: occupancy stays at 98% while debug-job p90 queue wait drops from 2 hours to 30 seconds

A write-up from real clusters on making it costly to game the scheduler.

Primary source Ai2 (Hugging Face Blog) · huggingface.co Read the original ↗

// Key points

  • Ai2's clusters range from 88 to 1024 GPUs and serve about 150 researchers, with demand running 2-3x supply.
  • Managers now set GPU time budgets enforced by a hierarchical fair-share scheduler: a declared minimum runtime is protected from preemption and charged to the budget, while unbudgeted time is free but preemptible.
  • Occupancy held at 98% before and after; over a 30-day test, teams received 98% of their owed GPU hours, with 13 of 15 allocations at 95% or more.
  • Debug-job p90 queue wait fell from 2 hours to 30 seconds, and human-in-the-loop repairs dropped by 74%.

Builder's takeThere's no released code, but the idea transfers directly: swap priority for budgets and grabbing resources gets a cost. When a team shares GPUs for inference or fine-tuning, I'd set budgets per project and keep a fast lane for debug jobs, which beats endlessly tuning priorities.