Circuit breaker
The "Broken Machine" Detector — stops sending traffic to a provider that has completely crashed.
problem it solves
Stops sending traffic to a provider that has completely crashed so a vendor outage does not break your app.
What it does
What it does: Stops sending traffic to a provider that has completely crashed.
What it watches: Hard errors (like a connection failing or the provider returning an error code).
When it triggers: If the provider fails a few times in a row (e.g., 5 straight errors).
The Action: It "trips" the circuit and temporarily stops sending any requests to that provider.
How it recovers: After a short timeout (like 10 seconds), it sends a single "test" request. If that succeeds, it turns the traffic back on.
What we need from you
- At least two providers onboarded for the models you callrequired
The breaker removes a provider from the candidate list. With one provider configured there is nowhere to route, so tripping turns an error into a different error. Add a second vendor key in Provider Vault before enabling this.
- Those providers must serve the same modelsrequired
Failover is per request, so the fallback has to be able to answer it. A second provider that carries none of the models you call is not a fallback.
- A week of shadow trafficrecommended
The defaults are chosen for general traffic, not yours. Shadow reports how often each provider would have been skipped, which is the only honest input to deciding whether the thresholds suit you.
- A provider preference, if you have onerecommended
Failover order is your routing order — cheapest-first, adjusted by any provider bias you have set. That is configured once in routing and applies to all traffic, not just failover. You do not configure it here.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Every provider stays in the candidate list however it is behaving, and a failing provider keeps receiving requests until something else stops it. | No circuit_breaker stage is recorded. |
| shadow | Nothing is routed differently. The breaker still tracks provider health and decides what it would have done, but every request is dispatched exactly as it would have been with the skill off. | A circuit_breaker stage appears with action=would_skip and the providers it would have removed — but only on requests where at least one circuit was open, so the stage's presence is the signal. |
| prod | Providers with an open circuit are dropped from the candidate list before dispatch, and traffic goes to the next feasible destination. If every candidate is open, all are kept and tried rather than failing the request outright. | Served destination reflects the failover; the skipped provider is absent from the attempt chain. |
Current policy
| What counts as an error | 429 or any 5xx | A 4xx other than 429 is a problem with the request, not the provider, and never trips the circuit. |
| Trip on consecutive failures | 5 in a row | Fires regardless of volume, so a hard outage opens the circuit quickly. |
| Trip on error rate | 15% over a 30s window | Only evaluated once at least 5 requests are in the window, so a quiet provider cannot trip on a single error. |
| Cooldown before retrying | 10s | After it elapses the circuit admits one probe request. A success closes it; a failure re-opens it. |
| Retry-After | Honoured as a hard block | A 429 carrying Retry-After blocks that provider until the deadline passes, independent of the circuit state. |
Worth knowing before you enable it
- ·Rate limits count as failures. A 429 means the provider is healthy and busy, not broken — but today it contributes to the error rate like a 500 does, so sustained throttling can open the circuit and send traffic to a more expensive destination.
- ·Failover is usually more expensive. Routing already picked the cheapest feasible destination, so the next one costs more by definition. The breaker trades money for availability, and that is the whole trade.
- ·Health is learned per gateway process. With several workers, each builds its own view of a provider, so they do not all route around it at the same moment.
- ·15% over 30 seconds is sensitive on low-volume traffic. One error in five recent requests is enough. This is the single most common reason to want different thresholds.
- ·This is the provider breaker, not the rollout breaker. A skill in `canary:N` or `prod` has a second, separate breaker on its scorecard — fail-open rate above 1.0%, turns per task up more than 25%, tool-call fidelity below 97%, task success or goodput more than 5 points under control — whose verdict is `CIRCUIT_BREAKER_TRIGGERED` on the scorecard and a `rolled_back` transition for the skill, not a skipped provider. Its thresholds and the lifecycle route are in the endpoint reference under **Skill lifecycle, canary and auto-revert**.
Difference between Circuit breaker, Adaptive concurrency, and Outlier ejection skills
| Circuit breaker | Adaptive concurrency | Outlier ejection | |
|---|---|---|---|
| the goal | Stop hitting a broken system. | Prevent a traffic jam. | Remove the worst server. |
| watches | Error responses. | Round-trip latency. | Success rate vs peers. |
| catches | A provider that is failing. | Slow but succeeding. | Worst of a healthy pool. |
| judged against | A fixed threshold. | Its own recent best. | Its peers' average. |
| fires when | 5 in a row, or 15% over 30s. | In-flight hits the limit. | 2 stdev below peers, or 3 5xx. |
| what happens to request | Next provider. | Next provider, or a 503. | Next server in the group. |
| can it error your caller | No. | Yes — 503 when all are capped. | No. |
| does it need backups? | Yes — nowhere else to route. | No. | Yes — no peers, no comparison. |
| recovery | 10s, then one probe. | Climbs back per healthy response. | 30s, straight back, no probe. |
What counts as a candidate
- Managed APIs (your keys)
- a provider in your Provider Vault — peers are the vault providers that serve that model — e.g. azure and openai for gpt-5.
- Self-hosted / GPU fleet
- a destination in your fleet — peers are the destinations carrying that model — two vLLM replicas, or two Azure regions.
What it replaces
- ·Retry wrappers around SDK calls, which retry the same failing provider rather than avoiding it.
- ·Hand-written try/except failover between two vendor SDKs, including the differing error shapes.
- ·Your own provider-health counters and rolling windows.
- ·Ad-hoc Retry-After parsing.
Full control over breaker policy is available on the enterprise tier.
- ·Dynamic thresholds — set your own error rate, window, sample floor and cooldown instead of the defaults above.
- ·Per-provider overrides — hold a flaky self-hosted endpoint to a different standard than a major vendor.
- ·Exponential backoff — increase the cooldown each time a provider re-opens, rather than retrying on a fixed interval.