Outlier ejection
The "Weakest Link" Remover — benches the worst-performing server in a group of identical servers.
problem it solves
Identifies and benches the worst-performing server in a group so a bad apple does not ruin your overall success rate.
What it does
What it does: Identifies and benches the worst-performing server in a group of identical servers.
What it watches: How one server's success rate compares to its peers.
When it triggers: When one specific server is performing noticeably worse than the rest of the group.
The Action: It temporarily kicks that specific "bad apple" out of the rotation so it stops ruining your overall success rate.
How it recovers: It puts the server in a "timeout" (e.g., 30 seconds), then quietly slips it back into the rotation to see if it has fixed itself.
What we need from you
- Two or more providers for the same modelrequired
On managed APIs that means two providers in your Provider Vault serving the same model. With one there are no peers to compare against, and the guard that never empties the candidate list puts it straight back — so the skill does nothing until a model has two.
- Enough traffic to comparerecommended
An endpoint needs 5 requests in the 60s window before it is scored.
- A week of shadow trafficrecommended
Shadow names the endpoints it would have ejected. Frequent ones on healthy traffic mean your pool is more uneven than the defaults assume.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Every endpoint stays in the pool however it compares. | No outlier_ejection stage is recorded. |
| shadow | Nothing is removed. Outcomes are still recorded and the comparison still runs, so the ejection it reports is real rather than hypothetical. | Stage with action=would_eject and the endpoints named, only when it would have ejected something. |
| prod | The request is still served — an ejected endpoint is removed from the candidate list, so it goes to the next one in your routing order instead. Nothing is skipped, delayed or rejected, and if every endpoint is ejected they are all kept rather than failing the request. This skill never returns an error to your caller. | Served destination reflects the ejection; the ejected endpoint is absent. |
Current policy
| Comparison window | 60s rolling | Judged on now, not on its worst hour. |
| Ejection threshold | 2 standard deviations below peer mean | Only when the spread is meaningful. |
| Minimum samples | 5 requests per endpoint | Two requests cannot establish a rate. |
| Consecutive 5xx | 3 in a row | A hard trip, no peer comparison needed. |
| Ejection duration | 30s, then back in the pool | It rejoins and is re-judged on its next requests. |
Worth knowing before you enable it
- ·A uniformly bad pool ejects nobody — nothing is an outlier. That case is the breaker's.
- ·Endpoints return after 30s with no probe; a still-broken one is ejected again within a few requests.
- ·Samples are per gateway process, so workers do not eject at the same moment.
- ·Any non-5xx counts as success, so a stream of 400s from a misconfigured endpoint reads as healthy.
Difference between Circuit breaker, Adaptive concurrency, and Outlier ejection skills
| Circuit breaker | Adaptive concurrency | Outlier ejection | |
|---|---|---|---|
| the goal | Stop hitting a broken system. | Prevent a traffic jam. | Remove the worst server. |
| watches | Error responses. | Round-trip latency. | Success rate vs peers. |
| catches | A provider that is failing. | Slow but succeeding. | Worst of a healthy pool. |
| judged against | A fixed threshold. | Its own recent best. | Its peers' average. |
| fires when | 5 in a row, or 15% over 30s. | In-flight hits the limit. | 2 stdev below peers, or 3 5xx. |
| what happens to request | Next provider. | Next provider, or a 503. | Next server in the group. |
| can it error your caller | No. | Yes — 503 when all are capped. | No. |
| does it need backups? | Yes — nowhere else to route. | No. | Yes — no peers, no comparison. |
| recovery | 10s, then one probe. | Climbs back per healthy response. | 30s, straight back, no probe. |
What counts as a candidate
- Managed APIs (your keys)
- a provider in your Provider Vault — peers are the vault providers that serve that model — e.g. azure and openai for gpt-5.
- Self-hosted / GPU fleet
- a destination in your fleet — peers are the destinations carrying that model — two vLLM replicas, or two Azure regions.
What it replaces
- ·Health checks that ask an endpoint if it is well rather than watching what it returns.
- ·Manual removal of a bad region or deployment from a rotation.
- ·Error dashboards watched by a human deciding when to pull one.
Full control over ejection policy is available on the enterprise tier.
- ·Custom thresholds, window and sample floor.
- ·Active probing before an ejected endpoint takes live traffic again.
- ·Per-pool policy, so a self-hosted pool is held to its own standard.