all docs

Spot / preemption reclaim

Monitors preemption webhooks and gracefully drains in-flight requests before node teardown.

problem it solves
Prevents request drops and errors when spot GPU instances are preempted by cloud providers.

What it does

What it does: Catches spot preemption notices and gracefully drains in-flight requests to warm replicas.

What it watches: Cloud provider preemption webhooks and 2-minute termination warnings.

When it triggers: Upon receiving a spot preemption warning signal from the cloud host.

The Action: Removes doomed node from candidate list and drains active streams to warm target nodes.

How it recovers: Requests complete without error while spot instance shuts down cleanly.

What we need from you

  • Heterogeneous cloud dispatch or backup poolrequired

    Requires active alternative nodes or clouds to receive drained traffic.

  • Spot instance preemption webhooksrequired

    Provider host must expose preemption termination notices.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Preempted spot nodes terminate abruptly, causing request errors.No spot_reclaim stage recorded.
shadowMonitors preemption signals and logs hypothetical drain timelines without re-routing.Stage with action=would_reclaim_spot logged.
prodExecutes zero-downtime node draining upon receiving preemption notice.Drained requests and 0-error spot termination logged.

Current policy

Notice window30s - 2min provider warning2 min on AWS/GCP; 30s on neoclouds (CoreWeave, Lambda).
Drain strategyStop new dispatch + graceful stream drainDrains active streams within provider notice window (30s-120s).
Reclaim error rateNear 0% for standard streamsStreams exceeding notice window require mid-stream state transfer.

Worth knowing before you enable it

  • ·Spot termination notices vary by provider (2 min on AWS/GCP, 30s on neoclouds like CoreWeave/Lambda).
  • ·Extremely long generation streams exceeding the provider warning window require state transfer.
  • ·Requires backup capacity available to absorb drained traffic.

What it replaces

  • ·Spot preemption 500 error cascades in production.
  • ·Custom node drain daemon scripts.
  • ·Fear of using ~70% discounted spot instances for inference.

Cross-provider emergency preemption drain and state checkpointing on enterprise.

  • ·Inter-cloud spot preemption traffic failover.
  • ·Mid-stream KV cache state checkpoint migration.
  • ·Spot market price-history predictive termination risk scoring.
team@acefleet.dev →