all docs

Utilization headroom governor

Maintains fleet load within target bands to ensure headroom for sudden bursts.

problem it solves
Prevents full GPU capacity saturation that leads to sudden latency spikes during load surges.

What it does

What it does: Holds fleet-wide GPU utilization within a target band (e.g. 65%-80%) over multi-minute windows.

What it watches: Fleet average GPU compute and memory utilization over rolling load history.

When it triggers: When fleet load drifts outside configured target headroom boundaries.

The Action: Adjusts autoscaler targets smoothly to ensure headroom for incoming traffic bursts.

How it recovers: Nudges capacity targets back down as sustained baseline traffic subsides.

What we need from you

  • GPU autoscaling skill enabledrequired

    Requires GPU autoscaling to execute capacity target adjustments.

  • At least 7 days of load historyrecommended

    Provides baseline data for historical headroom modeling.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Autoscaler operates purely on short-term reactive metrics.No utilization_headroom stage recorded.
shadowCalculates headroom targets and logs recommended governor adjustments.Stage with action=would_adjust_headroom logged.
prodApplies smooth target adjustments to autoscaler control loops.Target band occupancy and headroom adjustments logged.

Current policy

Target load band65% - 80% fleet utilizationOptimal balance of cost and burst safety.
Adjustment window5-minute smoothed moving averagePrevents short-term metric noise.
Safety buffer20% immediate headroom floorGuarantees capacity for sudden spikes.

Worth knowing before you enable it

  • ·Running at lower target utilization (e.g. 60%) increases standing infrastructure cost for higher availability.
  • ·Governor updates occur over minutes to avoid thrashing autoscaler provisioning.
  • ·Requires historical load data to optimize headroom bands accurately.

What it replaces

  • ·Manual autoscaler threshold tuning during high-traffic events.
  • ·Alert-fatigue from reactive capacity overload pagers.
  • ·Over-provisioning GPUs to 40% utilization out of fear.

Custom load band policies and event-aware headroom overrides on enterprise.

  • ·Scheduled headroom boosts for planned product launches.
  • ·Per-model class utilization target bands.
  • ·Cost-constrained headroom optimization modes.
team@acefleet.dev →