all docs

Storm guards

Rate-limits fleet scale-out and preemption drains to prevent stampede cascades.

problem it solves
Stops provider outage failovers and client bursts from triggering thundering-herd fleet collapse.

What it does

What it does: Rate-limits scale-out events and preemption drains to prevent fleet thundering herds.

What it watches: Concurrent node provisioning requests, failover burst rates, and control plane load.

When it triggers: During massive cloud outages or sudden 10x traffic surge stampedes.

The Action: Token-buckets autoscaling and drain events, queuing requests to maintain fleet stability.

How it recovers: Smooths out provisioning rate as traffic stabilizes back to steady-state.

What we need from you

Setup required before this can be enabled

Needs a fleet whose capacity is known — destinations declaring a ceiling.

  1. 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
  2. 2.A capacity ceiling on at least one destination — measure it with a benchmark, or declare the tokens/sec you provisioned. The floor is a fraction of known capacity, so with none declared there is nothing to take a fraction OF and the guard cannot bind.
Fleet Registration →
  • GPU autoscaling skill enabledrequired

    Governs autoscaler scale-out and preemption drain rate limits.

  • ACE control plane connectionrecommended

    Coordinates global token bucket limits across nodes.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Fleet responds to surges with unthrottled scale-out stampedes.No storm_guards stage recorded.
shadowMonitors failover rates and logs potential stampede conditions without throttling.Stage with action=would_guard_storm logged.
prodEnforces rate-limits on fleet scaling and drain actions under storm load.Throttled provision events and fleet stability metrics logged.

Current policy

Provisioning token bucketMax 5 nodes/min per classPrevents cloud API rate-limiting.
Global circuit breakerTied to control plane healthTrips if control plane is overloaded.
Stampede protectionCoordinated backoff queuesPrevents simultaneous cold-starts.

Worth knowing before you enable it

  • ·During sudden 100x bursts, storm guards queue new node provisioning to protect existing running nodes.
  • ·Cloud provider infrastructure APIs impose strict rate limits that storm guards protect against.
  • ·Requires GPU autoscaling to be enabled.

What it replaces

  • ·Thundering-herd cloud API rate limit blocks.
  • ·Fleet-wide crash cascades during cloud provider outages.
  • ·Uncontrolled cloud infrastructure bill spikes from runaway autoscaling.

Custom storm guard topologies and enterprise incident response rules on enterprise.

  • ·Custom token bucket refill rates per node class.
  • ·Cross-region storm isolation boundaries.
  • ·Integration with enterprise incident command webhooks.
team@acefleet.dev →