all docs

Circuit breaker

The "Broken Machine" Detector — stops sending traffic to a provider that has completely crashed.

problem it solves
Stops sending traffic to a provider that has completely crashed so a vendor outage does not break your app.

what is the difference between the 3 resilience skills? ↓

What it does

What it does: Stops sending traffic to a provider that has completely crashed.

What it watches: Hard errors (like a connection failing or the provider returning an error code).

When it triggers: If the provider fails a few times in a row (e.g., 5 straight errors).

The Action: It "trips" the circuit and temporarily stops sending any requests to that provider.

How it recovers: After a short timeout (like 10 seconds), it sends a single "test" request. If that succeeds, it turns the traffic back on.

What we need from you

  • At least two providers onboarded for the models you callrequired

    The breaker removes a provider from the candidate list. With one provider configured there is nowhere to route, so tripping turns an error into a different error. Add a second vendor key in Provider Vault before enabling this.

  • Those providers must serve the same modelsrequired

    Failover is per request, so the fallback has to be able to answer it. A second provider that carries none of the models you call is not a fallback.

  • A week of shadow trafficrecommended

    The defaults are chosen for general traffic, not yours. Shadow reports how often each provider would have been skipped, which is the only honest input to deciding whether the thresholds suit you.

  • A provider preference, if you have onerecommended

    Failover order is your routing order — cheapest-first, adjusted by any provider bias you have set. That is configured once in routing and applies to all traffic, not just failover. You do not configure it here.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Every provider stays in the candidate list however it is behaving, and a failing provider keeps receiving requests until something else stops it.No circuit_breaker stage is recorded.
shadowNothing is routed differently. The breaker still tracks provider health and decides what it would have done, but every request is dispatched exactly as it would have been with the skill off.A circuit_breaker stage appears with action=would_skip and the providers it would have removed — but only on requests where at least one circuit was open, so the stage's presence is the signal.
prodProviders with an open circuit are dropped from the candidate list before dispatch, and traffic goes to the next feasible destination. If every candidate is open, all are kept and tried rather than failing the request outright.Served destination reflects the failover; the skipped provider is absent from the attempt chain.

Current policy

What counts as an error429 or any 5xxA 4xx other than 429 is a problem with the request, not the provider, and never trips the circuit.
Trip on consecutive failures5 in a rowFires regardless of volume, so a hard outage opens the circuit quickly.
Trip on error rate15% over a 30s windowOnly evaluated once at least 5 requests are in the window, so a quiet provider cannot trip on a single error.
Cooldown before retrying10sAfter it elapses the circuit admits one probe request. A success closes it; a failure re-opens it.
Retry-AfterHonoured as a hard blockA 429 carrying Retry-After blocks that provider until the deadline passes, independent of the circuit state.

Worth knowing before you enable it

  • ·Rate limits count as failures. A 429 means the provider is healthy and busy, not broken — but today it contributes to the error rate like a 500 does, so sustained throttling can open the circuit and send traffic to a more expensive destination.
  • ·Failover is usually more expensive. Routing already picked the cheapest feasible destination, so the next one costs more by definition. The breaker trades money for availability, and that is the whole trade.
  • ·Health is learned per gateway process. With several workers, each builds its own view of a provider, so they do not all route around it at the same moment.
  • ·15% over 30 seconds is sensitive on low-volume traffic. One error in five recent requests is enough. This is the single most common reason to want different thresholds.
  • ·This is the provider breaker, not the rollout breaker. A skill in `canary:N` or `prod` has a second, separate breaker on its scorecard — fail-open rate above 1.0%, turns per task up more than 25%, tool-call fidelity below 97%, task success or goodput more than 5 points under control — whose verdict is `CIRCUIT_BREAKER_TRIGGERED` on the scorecard and a `rolled_back` transition for the skill, not a skipped provider. Its thresholds and the lifecycle route are in the endpoint reference under **Skill lifecycle, canary and auto-revert**.

Difference between Circuit breaker, Adaptive concurrency, and Outlier ejection skills

Circuit breakerAdaptive concurrencyOutlier ejection
the goalStop hitting a broken system.Prevent a traffic jam.Remove the worst server.
watchesError responses.Round-trip latency.Success rate vs peers.
catchesA provider that is failing.Slow but succeeding.Worst of a healthy pool.
judged againstA fixed threshold.Its own recent best.Its peers' average.
fires when5 in a row, or 15% over 30s.In-flight hits the limit.2 stdev below peers, or 3 5xx.
what happens to requestNext provider.Next provider, or a 503.Next server in the group.
can it error your callerNo.Yes — 503 when all are capped.No.
does it need backups?Yes — nowhere else to route.No.Yes — no peers, no comparison.
recovery10s, then one probe.Climbs back per healthy response.30s, straight back, no probe.

What counts as a candidate

Managed APIs (your keys)
a provider in your Provider Vault — peers are the vault providers that serve that model — e.g. azure and openai for gpt-5.
Self-hosted / GPU fleet
a destination in your fleet — peers are the destinations carrying that model — two vLLM replicas, or two Azure regions.

What it replaces

  • ·Retry wrappers around SDK calls, which retry the same failing provider rather than avoiding it.
  • ·Hand-written try/except failover between two vendor SDKs, including the differing error shapes.
  • ·Your own provider-health counters and rolling windows.
  • ·Ad-hoc Retry-After parsing.

Full control over breaker policy is available on the enterprise tier.

  • ·Dynamic thresholds — set your own error rate, window, sample floor and cooldown instead of the defaults above.
  • ·Per-provider overrides — hold a flaky self-hosted endpoint to a different standard than a major vendor.
  • ·Exponential backoff — increase the cooldown each time a provider re-opens, rather than retrying on a fixed interval.
team@acefleet.dev →