all docs
/ errors · HTTP 503

overloaded

Gateway capacity shed: ingress budget, loop stall, queue wait, or dispatch limit.

What it means

ACE shed the request on its own capacity, never the provider's. Three places decide it: the transport, before the body is read (declared size against what is in flight, ACE_INGRESS_BUDGET, off unless set; or a heavy body while the event loop has been stalled past ACE_LOOP_STALL_SHED_SHARE of the last ACE_LOOP_STALL_WINDOW_S, on by default at 0.5 over 10s — light bodies are still served); engine entry, when the request sat past ACE_QUEUE_WAIT_MAX_S (10s) before the engine reached it; and dispatch, when every destination was at its concurrency limit. Always carries Retry-After, computed from what this process has actually drained lately (1–30s). The body wears the surface's own overload envelope (overloaded_error on the Anthropic surfaces). Back off and retry, or send direct to the provider — neither the key nor the provider is involved. GET /healthz reports the gauge under load.

How to recognise it

The gateway answers HTTP 503 with x-ace-error: overloaded on the response, and a body in the shared error envelope:

{
  "error": {
    "message": "...",
    "type": "overloaded",
    "code": 503
  }
}

On a vendor-shaped surface (/anthropic/…, /gemini/…, /bedrock/…) the body wears that vendor's error envelope instead; the x-ace-error header is the same everywhere.

Is it ACE or the provider?

x-ace-error is present only on errors ACE originated. A vendor error relayed from upstream — a real provider 429, a provider 401 for a bad pass-through key — carries no x-ace-error, and its own type and code mean what the provider says. Read the header before deciding whether to retry, re-mint a key or surface the error.