Storm guards
Rate-limits fleet scale-out and preemption drains to prevent stampede cascades.
problem it solves
Stops provider outage failovers and client bursts from triggering thundering-herd fleet collapse.
What it does
What it does: Rate-limits scale-out events and preemption drains to prevent fleet thundering herds.
What it watches: Concurrent node provisioning requests, failover burst rates, and control plane load.
When it triggers: During massive cloud outages or sudden 10x traffic surge stampedes.
The Action: Token-buckets autoscaling and drain events, queuing requests to maintain fleet stability.
How it recovers: Smooths out provisioning rate as traffic stabilizes back to steady-state.
What we need from you
Setup required before this can be enabled
Needs a fleet whose capacity is known — destinations declaring a ceiling.
- 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
- 2.A capacity ceiling on at least one destination — measure it with a benchmark, or declare the tokens/sec you provisioned. The floor is a fraction of known capacity, so with none declared there is nothing to take a fraction OF and the guard cannot bind.
- GPU autoscaling skill enabledrequired
Governs autoscaler scale-out and preemption drain rate limits.
- ACE control plane connectionrecommended
Coordinates global token bucket limits across nodes.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Fleet responds to surges with unthrottled scale-out stampedes. | No storm_guards stage recorded. |
| shadow | Monitors failover rates and logs potential stampede conditions without throttling. | Stage with action=would_guard_storm logged. |
| prod | Enforces rate-limits on fleet scaling and drain actions under storm load. | Throttled provision events and fleet stability metrics logged. |
Current policy
| Provisioning token bucket | Max 5 nodes/min per class | Prevents cloud API rate-limiting. |
| Global circuit breaker | Tied to control plane health | Trips if control plane is overloaded. |
| Stampede protection | Coordinated backoff queues | Prevents simultaneous cold-starts. |
Worth knowing before you enable it
- ·During sudden 100x bursts, storm guards queue new node provisioning to protect existing running nodes.
- ·Cloud provider infrastructure APIs impose strict rate limits that storm guards protect against.
- ·Requires GPU autoscaling to be enabled.
What it replaces
- ·Thundering-herd cloud API rate limit blocks.
- ·Fleet-wide crash cascades during cloud provider outages.
- ·Uncontrolled cloud infrastructure bill spikes from runaway autoscaling.
Custom storm guard topologies and enterprise incident response rules on enterprise.
- ·Custom token bucket refill rates per node class.
- ·Cross-region storm isolation boundaries.
- ·Integration with enterprise incident command webhooks.