all docs

GPU autoscaling

Queue-depth and latency-aware capacity scale-up demand signaling with costed recommendations.

problem it solves
Prevents unplaceable workloads from silently starving in pending queues while eliminating over-provisioned idle capacity.

What it does

What it does: Evaluates unsatisfied queue depth and TTFT latency pressure across unplaceable pods to emit explicit scale-up recommendations.

What it watches: Unplaced pending pod queue depth, target TTFT budgets, current cluster VRAM headroom, and cold-start latency.

When it triggers: When pending workloads cannot be placed into existing GPU slices or nodes within target latency budgets.

The Action: Emits explicit, costed scale-up requests consumable by external provisioners (Karpenter / KEDA / Cluster-Autoscaler).

How it recovers: Recommends scaling down idle replicas smoothly after traffic bursts decay.

What we need from you

  • Registered GPU fleetrequired

    Requires cloud or on-prem GPU infrastructure connected to ACE control plane.

  • ACE node agent installedrequired

    Reports real-time queue depth, placement failures, and latency signals.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Pending unplaced workloads queue silently without scale-up evaluation.No gpu_autoscaling stage recorded.
shadowEvaluates unplaced pod queue depth and logs costed scale-up recommendations without triggering alerts.Stage with action=would_scale_gpu and recommended_replicas logged.
prodPublishes structured, costed scale-up requests to external infrastructure provisioners and autoscaler control loops.Scale-up demand requests and shortfall metrics logged.

Current policy

Scaling signalQueue depth + p95 latency shortfallDirectly measures unfulfilled demand rather than trailing GPU utilization.
Evaluation window30s smoothing windowFilters transient load spikes before emitting scale-up requests.
Target provisionerExternal provisioner (Karpenter / KEDA / Cluster-Autoscaler)Integrates with standard cloud-native infrastructure controllers.

Worth knowing before you enable it

  • ·GPU node cold-start times (container pull + weight loading) take 1-3 minutes without warm pooling.
  • ·Scaling signals based on GPU VRAM alone are misleading due to KV cache allocation.
  • ·Cloud provider quota limits cap maximum scale-out replica bounds.

What it replaces

  • ·Naive CPU/Memory-based Kubernetes HPA setups.
  • ·Silent pending pod starvation in overloaded clusters.
  • ·Over-provisioned 24/7 idle GPU clusters.

Cross-cloud multi-region autoscaling and pre-warm scheduling on enterprise.

  • ·Predictive calendar and traffic-based pre-warming.
  • ·Multi-cloud provider spot/on-demand failover scaling.
  • ·Custom scaling metric formula builders.
team@acefleet.dev →