overloaded
Gateway capacity shed: ingress budget, loop stall, queue wait, or dispatch limit.
What it means
ACE shed the request on its own capacity, never the provider's. Three places decide it: the transport, before the body is read (declared size against what is in flight, ACE_INGRESS_BUDGET, off unless set; or a heavy body while the event loop has been stalled past ACE_LOOP_STALL_SHED_SHARE of the last ACE_LOOP_STALL_WINDOW_S, on by default at 0.5 over 10s — light bodies are still served); engine entry, when the request sat past ACE_QUEUE_WAIT_MAX_S (10s) before the engine reached it; and dispatch, when every destination was at its concurrency limit. Always carries Retry-After, computed from what this process has actually drained lately (1–30s). The body wears the surface's own overload envelope (overloaded_error on the Anthropic surfaces). Back off and retry, or send direct to the provider — neither the key nor the provider is involved. GET /healthz reports the gauge under load.
How to recognise it
The gateway answers HTTP 503 with x-ace-error: overloaded on the response, and a body in the shared error envelope:
{
"error": {
"message": "...",
"type": "overloaded",
"code": 503
}
}
On a vendor-shaped surface (/anthropic/…, /gemini/…, /bedrock/…) the body wears that vendor's error envelope instead; the x-ace-error header is the same everywhere.
Is it ACE or the provider?
x-ace-error is present only on errors ACE originated. A vendor error relayed from upstream — a real provider 429, a provider 401 for a bad pass-through key — carries no x-ace-error, and its own type and code mean what the provider says. Read the header before deciding whether to retry, re-mint a key or surface the error.
Related
- Error contracts & status codes: the full taxonomy
- Response header specification
- Echo mode: reproduce a call without spending