Adaptive concurrency
The "Traffic Jam" Preventer — stops a sluggish provider from causing a system-wide backlog.
problem it solves
Stops a sluggish provider from causing a system-wide backlog when slow responses tie up resources.
What it does
What it does: Stops a sluggish provider from causing a system-wide backlog.
What it watches: Speed (how long it takes the provider to respond).
When it triggers: When the provider starts taking longer than its own recent normal speed.
The Action: It lowers the limit on how many requests that provider is allowed to handle at once. Any extra requests are instantly handed to your next backup provider.
How it recovers: As the provider speeds back up, the system gradually raises its capacity limit again.
What we need from you
- Nothing to configurerequired
It learns each destination's normal latency from its own traffic. No baseline to supply, no per-model setup.
- A second providerrecommended
Not required — the cap protects your other models either way. A second provider in your Provider Vault is what lets shed requests be served instead of rejected. With one, expect some 503s that would previously have succeeded slowly.
- Steady trafficrecommended
The baseline is a 60s rolling window, so a destination seeing a few requests an hour never moves far from its default.
- A week of shadow trafficrecommended
Shadow reports how often each destination would have shed. Near zero on healthy traffic; if not, prod would reject real requests.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Any number of requests may be in flight to a destination. | No adaptive_concurrency stage is recorded. |
| shadow | Nothing is shed. It still takes slots and measures RTT so the limit keeps moving, but every request is dispatched as it would have been with the skill off. | Stage with action=would_shed, only on requests where something would have been shed — so its presence is the signal. |
| prod | A destination at its limit is skipped for the next candidate. With every candidate at its limit, the request is shed with 503 and no upstream call is placed. | Stage with action=shed. The served destination reflects the fall-forward. |
Current policy
| Starting limit | 20 in flight per destination | Before it has latency history. |
| Range | 5 to 200 | Never zero, never unbounded. |
| Latency baseline | fastest RTT in the last 60s | Rolling, so one fast response cannot define 'normal' forever. |
| On a slow response | limit x (fastest / current), floored at half | Smoothed; one slow response moves it very little. |
| On an error | limit x 0.8, immediately | Harder evidence than latency. |
| Recovery | +sqrt(limit) per healthy response | Climbs back rather than staying where its worst moment left it. |
Worth knowing before you enable it
- ·Shedding is not failover. With every provider at its limit the caller gets a 503, which your code has to handle.
- ·Limits are per gateway process, so the effective ceiling is the limit times your worker count.
- ·It reacts to latency only. A provider returning fast wrong answers looks healthy.
- ·A large and a small model on one destination share a baseline, so the large one's normal latency reads as congestion.
- ·Shadow reports would_shed on traffic that was served fine. That is the point of it.
Difference between Circuit breaker, Adaptive concurrency, and Outlier ejection skills
| Circuit breaker | Adaptive concurrency | Outlier ejection | |
|---|---|---|---|
| the goal | Stop hitting a broken system. | Prevent a traffic jam. | Remove the worst server. |
| watches | Error responses. | Round-trip latency. | Success rate vs peers. |
| catches | A provider that is failing. | Slow but succeeding. | Worst of a healthy pool. |
| judged against | A fixed threshold. | Its own recent best. | Its peers' average. |
| fires when | 5 in a row, or 15% over 30s. | In-flight hits the limit. | 2 stdev below peers, or 3 5xx. |
| what happens to request | Next provider. | Next provider, or a 503. | Next server in the group. |
| can it error your caller | No. | Yes — 503 when all are capped. | No. |
| does it need backups? | Yes — nowhere else to route. | No. | Yes — no peers, no comparison. |
| recovery | 10s, then one probe. | Climbs back per healthy response. | 30s, straight back, no probe. |
What counts as a candidate
- Managed APIs (your keys)
- a provider in your Provider Vault — peers are the vault providers that serve that model — e.g. azure and openai for gpt-5.
- Self-hosted / GPU fleet
- a destination in your fleet — peers are the destinations carrying that model — two vLLM replicas, or two Azure regions.
What it replaces
- ·Static per-provider concurrency caps.
- ·Hand-tuned semaphores and connection-pool limits.
- ·Client-side timeouts used as backpressure.
- ·p99 watchdogs and the runbook step that says 'turn down concurrency'.
Full control over the limiter is available on the enterprise tier.
- ·Custom bands per destination.
- ·Per-model baselines, so a reasoning model and a small one do not share one.
- ·A bounded queue admitted in order, instead of shedding.