all docs

Adaptive concurrency

The "Traffic Jam" Preventer — stops a sluggish provider from causing a system-wide backlog.

problem it solves
Stops a sluggish provider from causing a system-wide backlog when slow responses tie up resources.

what is the difference between the 3 resilience skills? ↓

What it does

What it does: Stops a sluggish provider from causing a system-wide backlog.

What it watches: Speed (how long it takes the provider to respond).

When it triggers: When the provider starts taking longer than its own recent normal speed.

The Action: It lowers the limit on how many requests that provider is allowed to handle at once. Any extra requests are instantly handed to your next backup provider.

How it recovers: As the provider speeds back up, the system gradually raises its capacity limit again.

What we need from you

  • Nothing to configurerequired

    It learns each destination's normal latency from its own traffic. No baseline to supply, no per-model setup.

  • A second providerrecommended

    Not required — the cap protects your other models either way. A second provider in your Provider Vault is what lets shed requests be served instead of rejected. With one, expect some 503s that would previously have succeeded slowly.

  • Steady trafficrecommended

    The baseline is a 60s rolling window, so a destination seeing a few requests an hour never moves far from its default.

  • A week of shadow trafficrecommended

    Shadow reports how often each destination would have shed. Near zero on healthy traffic; if not, prod would reject real requests.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Any number of requests may be in flight to a destination.No adaptive_concurrency stage is recorded.
shadowNothing is shed. It still takes slots and measures RTT so the limit keeps moving, but every request is dispatched as it would have been with the skill off.Stage with action=would_shed, only on requests where something would have been shed — so its presence is the signal.
prodA destination at its limit is skipped for the next candidate. With every candidate at its limit, the request is shed with 503 and no upstream call is placed.Stage with action=shed. The served destination reflects the fall-forward.

Current policy

Starting limit20 in flight per destinationBefore it has latency history.
Range5 to 200Never zero, never unbounded.
Latency baselinefastest RTT in the last 60sRolling, so one fast response cannot define 'normal' forever.
On a slow responselimit x (fastest / current), floored at halfSmoothed; one slow response moves it very little.
On an errorlimit x 0.8, immediatelyHarder evidence than latency.
Recovery+sqrt(limit) per healthy responseClimbs back rather than staying where its worst moment left it.

Worth knowing before you enable it

  • ·Shedding is not failover. With every provider at its limit the caller gets a 503, which your code has to handle.
  • ·Limits are per gateway process, so the effective ceiling is the limit times your worker count.
  • ·It reacts to latency only. A provider returning fast wrong answers looks healthy.
  • ·A large and a small model on one destination share a baseline, so the large one's normal latency reads as congestion.
  • ·Shadow reports would_shed on traffic that was served fine. That is the point of it.

Difference between Circuit breaker, Adaptive concurrency, and Outlier ejection skills

Circuit breakerAdaptive concurrencyOutlier ejection
the goalStop hitting a broken system.Prevent a traffic jam.Remove the worst server.
watchesError responses.Round-trip latency.Success rate vs peers.
catchesA provider that is failing.Slow but succeeding.Worst of a healthy pool.
judged againstA fixed threshold.Its own recent best.Its peers' average.
fires when5 in a row, or 15% over 30s.In-flight hits the limit.2 stdev below peers, or 3 5xx.
what happens to requestNext provider.Next provider, or a 503.Next server in the group.
can it error your callerNo.Yes — 503 when all are capped.No.
does it need backups?Yes — nowhere else to route.No.Yes — no peers, no comparison.
recovery10s, then one probe.Climbs back per healthy response.30s, straight back, no probe.

What counts as a candidate

Managed APIs (your keys)
a provider in your Provider Vault — peers are the vault providers that serve that model — e.g. azure and openai for gpt-5.
Self-hosted / GPU fleet
a destination in your fleet — peers are the destinations carrying that model — two vLLM replicas, or two Azure regions.

What it replaces

  • ·Static per-provider concurrency caps.
  • ·Hand-tuned semaphores and connection-pool limits.
  • ·Client-side timeouts used as backpressure.
  • ·p99 watchdogs and the runbook step that says 'turn down concurrency'.

Full control over the limiter is available on the enterprise tier.

  • ·Custom bands per destination.
  • ·Per-model baselines, so a reasoning model and a small one do not share one.
  • ·A bounded queue admitted in order, instead of shedding.
team@acefleet.dev →