Spot / preemption reclaim
Monitors preemption webhooks and gracefully drains in-flight requests before node teardown.
problem it solves
Prevents request drops and errors when spot GPU instances are preempted by cloud providers.
What it does
What it does: Catches spot preemption notices and gracefully drains in-flight requests to warm replicas.
What it watches: Cloud provider preemption webhooks and 2-minute termination warnings.
When it triggers: Upon receiving a spot preemption warning signal from the cloud host.
The Action: Removes doomed node from candidate list and drains active streams to warm target nodes.
How it recovers: Requests complete without error while spot instance shuts down cleanly.
What we need from you
- Heterogeneous cloud dispatch or backup poolrequired
Requires active alternative nodes or clouds to receive drained traffic.
- Spot instance preemption webhooksrequired
Provider host must expose preemption termination notices.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Preempted spot nodes terminate abruptly, causing request errors. | No spot_reclaim stage recorded. |
| shadow | Monitors preemption signals and logs hypothetical drain timelines without re-routing. | Stage with action=would_reclaim_spot logged. |
| prod | Executes zero-downtime node draining upon receiving preemption notice. | Drained requests and 0-error spot termination logged. |
Current policy
| Notice window | 30s - 2min provider warning | 2 min on AWS/GCP; 30s on neoclouds (CoreWeave, Lambda). |
| Drain strategy | Stop new dispatch + graceful stream drain | Drains active streams within provider notice window (30s-120s). |
| Reclaim error rate | Near 0% for standard streams | Streams exceeding notice window require mid-stream state transfer. |
Worth knowing before you enable it
- ·Spot termination notices vary by provider (2 min on AWS/GCP, 30s on neoclouds like CoreWeave/Lambda).
- ·Extremely long generation streams exceeding the provider warning window require state transfer.
- ·Requires backup capacity available to absorb drained traffic.
What it replaces
- ·Spot preemption 500 error cascades in production.
- ·Custom node drain daemon scripts.
- ·Fear of using ~70% discounted spot instances for inference.
Cross-provider emergency preemption drain and state checkpointing on enterprise.
- ·Inter-cloud spot preemption traffic failover.
- ·Mid-stream KV cache state checkpoint migration.
- ·Spot market price-history predictive termination risk scoring.