GPU autoscaling
Queue-depth and latency-aware capacity scale-up demand signaling with costed recommendations.
problem it solves
Prevents unplaceable workloads from silently starving in pending queues while eliminating over-provisioned idle capacity.
What it does
What it does: Evaluates unsatisfied queue depth and TTFT latency pressure across unplaceable pods to emit explicit scale-up recommendations.
What it watches: Unplaced pending pod queue depth, target TTFT budgets, current cluster VRAM headroom, and cold-start latency.
When it triggers: When pending workloads cannot be placed into existing GPU slices or nodes within target latency budgets.
The Action: Emits explicit, costed scale-up requests consumable by external provisioners (Karpenter / KEDA / Cluster-Autoscaler).
How it recovers: Recommends scaling down idle replicas smoothly after traffic bursts decay.
What we need from you
- Registered GPU fleetrequired
Requires cloud or on-prem GPU infrastructure connected to ACE control plane.
- ACE node agent installedrequired
Reports real-time queue depth, placement failures, and latency signals.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Pending unplaced workloads queue silently without scale-up evaluation. | No gpu_autoscaling stage recorded. |
| shadow | Evaluates unplaced pod queue depth and logs costed scale-up recommendations without triggering alerts. | Stage with action=would_scale_gpu and recommended_replicas logged. |
| prod | Publishes structured, costed scale-up requests to external infrastructure provisioners and autoscaler control loops. | Scale-up demand requests and shortfall metrics logged. |
Current policy
| Scaling signal | Queue depth + p95 latency shortfall | Directly measures unfulfilled demand rather than trailing GPU utilization. |
| Evaluation window | 30s smoothing window | Filters transient load spikes before emitting scale-up requests. |
| Target provisioner | External provisioner (Karpenter / KEDA / Cluster-Autoscaler) | Integrates with standard cloud-native infrastructure controllers. |
Worth knowing before you enable it
- ·GPU node cold-start times (container pull + weight loading) take 1-3 minutes without warm pooling.
- ·Scaling signals based on GPU VRAM alone are misleading due to KV cache allocation.
- ·Cloud provider quota limits cap maximum scale-out replica bounds.
What it replaces
- ·Naive CPU/Memory-based Kubernetes HPA setups.
- ·Silent pending pod starvation in overloaded clusters.
- ·Over-provisioned 24/7 idle GPU clusters.
Cross-cloud multi-region autoscaling and pre-warm scheduling on enterprise.
- ·Predictive calendar and traffic-based pre-warming.
- ·Multi-cloud provider spot/on-demand failover scaling.
- ·Custom scaling metric formula builders.