all docs
/ reference

Response Header Specification

What the gateway reports back on the OpenAI-shaped surface: cache result, chosen route, cost, budget, x-ace-error on an ACE-originated error, x-ace-skill-modes and x-ace-skills-applied, the vendor's own rate-limit headers relayed verbatim, PII redaction counts and stages (and the budget cut), the prompt- and trajectory-compaction verdicts with their token counts and checkpoint digest, and the skill latency ceiling behind ceiling_exceeded, overloaded and loop_stalled.

The response body is byte-shape unmodified — a standard OpenAI chat.completion object — so existing SDKs parse it with zero changes. Everything ACE-specific rides on headers.

Header Meaning
x-ace-cache Semantic cache outcome: HIT, MISS, or BYPASS (-similarity, -threshold, -miss-reason).
MoreHIT / MISS / BYPASS — semantic cache result. BYPASS says the lookup was not allowed to serve and x-ace-cache-bypass-reason names the gate that refused it: skill_off, shadow_mode, no_cache_requested (execution.no_cache), cache_control_no_cache (the header), high_temperature (temperature > 0.8), stateful_unpartitioned (the strict key policy refusing an unpartitioned stateful request), provider_evaluated_extension (a guardrailConfig / safetySettings that must run at the provider), or tool_result_turn — the live turn is a tool result, an agent step. Under the read_set and strict key policies that step's key carries its own tool-result digest and the session, so nothing but a byte-identical re-send can match it, and that re-send is served from the exact-prompt table without an embed: a client retry is still a HIT at ~0ms, and no MISS is reported because the embed and vector query that would have produced it never run — nor the shadow lookup, nor the vector write. skill_params.semantic_cache.lookup_on_tool_results: true restores the approximate tier on such steps; under the text key policy it runs regardless. x-ace-cache-shadow carries what the lookup would have done (WOULD_HIT / MISS / PENDING — PENDING means the shadow lookup had not answered within the join grace, so count it as unmeasured, never as a miss). With it: -similarity (how close the best candidate came, on a miss too), -threshold (the floor this lookup actually applied — adaptive thresholds raise it to 0.98 on an entity-dense prompt, so it is not always the cache's base setting) and, on a MISS, -miss-reason: no_candidate (nothing stored in this namespace yet), below_threshold, identifier_mismatch (the numeric guard: an identifier token differs between the two prompts, so a 0.98 that missed is the guard refusing a different thing, not the cache failing. Four shapes make a token an identifier — it carries a digit (a ticket number, a hex id, a version); it carries an underscore, which ordinary English does not (test_staging vs test_production); it is upper-case and written after ticket / case / order / id / …; or it is written after a word naming a SLOT rather than a code — database, user, branch, service, instance, cluster, … — where the next token is the name whatever its case, which is what separates user alice from user bob and service payment-gateway from service auth-gateway. A bare hyphen does NOT make an identifier, so read-only still matches read only), polarity_mismatch (the two prompts name OPPOSITE sides of one action while both stay affirmative — allow vs block, enable vs disable, grant vs revoke — a pair the embedder puts ~0.98 apart because every other word matches. It fires only when each side carries one pole and not the other, so a sentence naming both, the right way to start and stop the cluster, is ordinary phrasing rather than a conflict), negation_mismatch (a negator present on one side only, or — both sides negating — one side's negators a proper subset of the other's, the extra negator stacked on a negation they share: is it NOT possible to avoid downtime against is it possible to avoid downtime. Two disjoint negators are left alone, being two ways to say one negative), word_order, store_error (the vector store or the embedder failed; the trace stage's store_error names the exception class), and the three the skill latency ceiling reports — ceiling_exceeded, overloaded, loop_stalled (the lookup did not answer inside its per-request ceiling and the request went upstream rather than wait; see The skill latency ceiling below).
x-ace-cache-usage / x-ace-cache-entry-age-s / x-ace-cache-key-policy Replayed body indicator, cache entry age in seconds, and active key policy.
MoreOn a HIT, x-ace-cache-usage: replayed says the body is the STORED completion served verbatim — its id and its usage are the FIRST call's. They are deliberately not rewritten (the surface contract is a vendor-shaped body, and callers parse it), so a hit can show a vendor cache_creation_input_tokens beside x-ace-cost-usd: 0.000000 on a request that never reached the vendor. Reconcile spend against x-ace-cost-usd, never a replayed usage, and drop replayed responses before measuring anything about the vendor's own prompt cache; the same header rides a completion replay (x-ace-replay: HIT). -entry-age-s is how old the served entry was. -key-policy is the key policy this lookup ran under (read_set is the default, with a session partition).
x-ace-cache-scope Cache isolation scope when partitioned (partition, session, isolated).
MorePresent only when this request did NOT share the org-wide cache pool: partition (a valid x-ace-cache-partition was sent), session (a deployment in session-isolation mode), isolated (fail-closed — alone in its namespace, always MISS), bypass-unpartitioned (the strict cache-key policy refused to serve an unpartitioned stateful request). Absent on an ordinary production request.
x-ace-cache-entity Digest of prompt identifier tokens for gateway entity partition isolation.
MoreA digest of the identifier tokens the gateway found in this prompt, present only when it found any and ACE_CACHE_ENTITY_PARTITION is on (default off). It names a namespace the GATEWAY imposed: two prompts whose identifier sets differ carry different digests and cannot read each other's entries at any similarity. It is NOT x-ace-cache-scope, which reports the isolation the CALLER asked for (session / partition / isolated); this one reports isolation applied on top of whatever the caller asked for, and both can appear on the same response. The partition is keyed on the same tokens identifier_mismatch compares, so it separates only the pairs that guard was already refusing — no hit served today is lost — and it reaches the ones the guard could not: the index is queried one candidate deep, so a correct entry sitting behind a same-shape-different-id neighbour was unreachable, and inside its own partition it is first. What the header adds is legibility. Without it, a probe that missed because it named a different entity reports no_candidate at similarity 0.0, which a reader cannot tell apart from a cold cache; with it, the digest says which namespace the lookup actually ran in. Turning it on re-namespaces every stored entry whose prompt contains an identifier token — a partial cold start, silent by construction. It is gated in lockstep with the identifier guard, so an operator who turned that guard off does not still get identifier separation imposed structurally.
x-ace-served-by Serving mechanism: echo, direct_passthrough, or upstream engine.
MoreDestination id that actually served this request. echo means no real upstream is configured; direct_passthrough means the optimization pipeline was bypassed.
x-ace-destination Specific upstream provider or fleet node that generated completion.
MoreWhich destination or provider actually answered: a fleet destination id on a dispatched request, the provider name (anthropic, azure, bedrock, …) on a relayed one, echo on an echo-served one. On the relay path x-ace-served-by names only the MECHANISM — every relayed response reports direct_passthrough whichever provider answered — so this is the field that tells two relayed responses apart.
x-ace-cost-usd Actual cost in USD (or estimated cost if streaming).
MoreActual (or, when streaming, estimated) cost of this request.
x-ace-route-model / x-ace-route-tier / x-ace-route-reason Selected model, routing tier (smart, fast, lite), and routing rationale.
MoreWhich model and tier the router picked, and why. Present when a router is configured. A router that ran and routed nothing says why in -reason: no_candidates, no_tier_choice, router_failed, or ceiling_exceeded / overloaded / loop_stalled when the skill latency ceiling shed the stage and the caller's model passed through. Under the agent trajectory router the tier is the policy alias (smart / fast / lite) and a reason of continuity:deferred:<why> means the pinned model served.
x-ace-route-escalated true if output failed validation and escalated to flagship model.
Moretrue if a cheap-tier answer failed validation and escalated to a flagship model.
x-ace-warnings Comma-separated codes for fallback/bypass warnings (see warning codes table).
MoreComma-separated codes, each defined in the table below: echo_fallback, model_auto_selected, provider_key_pass_through, provider_key_ignored, extensions_omitted, optimization_suppressed, pipeline_bypassed, agent_trajectory_router_reverted. Absent when there is nothing to warn about — an ordinary relay is not a warning.
x-ace-budget-remaining-daily Remaining daily team budget when a finite limit is configured.
MoreRemaining daily team budget, when a finite cap exists.
x-ace-error Standardized ACE taxonomy error type on gateway-originated errors.
MoreOn every error ACE originated, on every surface: the ACE taxonomy type (invalid_developer_key, budget_exhausted, request_error, skill_locked, …) whatever envelope the body wears. Absent on a vendor error relayed from upstream — that absence is the signal a retry classifier branches on. Values in the errors section.
x-ace-skills-applied Echoes the skill_id=mode overrides that actually ran on this request.
MoreThe skill_id=mode pairs from the request's x-ace-skills header that this request actually ran with (semantic_cache=off,prompt_compaction=shadow). Only the pairs that applied: one the body's extra_body.ace.skill_overrides beat, or that compute_status: passthrough_only forced to shadow after it was accepted, is absent — as is the header when nothing was sent or nothing applied. The per-skill headers (x-ace-cache, x-ace-compaction-*, x-ace-skill-modes) report exactly as before. A * wildcard is echoed expanded — one pair per registered universal skill it resolved to, never * itself — and a locked skill a wildcard skipped is reported as <skill>=locked, which is not a mode, so a job asserting pii_ner=off fails on it rather than reading a lock as an off.
x-ace-skill-modes Sorted list of all active skills and effective execution modes (prod/shadow).
MoreEvery skill that ran on this request and the mode it ran under, as sorted skill_id=mode pairs (llm_router=shadow,semantic_cache=prod) — whatever set the mode: the key's stored lifecycle, the deployment default or a per-request override. One header rather than a flag per skill because the question is asked about the request as a whole: which of these numbers am I allowed to believe changed something. A skill that was off is absent, and so is a retired skill id still stored on the key with no module behind it. Distinct from x-ace-skills-applied, which echoes only the pairs you sent in x-ace-skills that took effect; the two never disagree, because both report the intent each skill ran under.
x-ace-knowledge-graph What skill_knowledge_graph added to the request, or why it added nothing.
MoreThe procedure-recall skill's outcome on this request. injected;tokens=<n>;source=<proc_id> — a stored procedure block was added to the request; would_inject;tokens=<n>;source=<proc_id> — the same in shadow mode, where nothing is added; skipped;reason=<budget_exceeded|cache_hit|replay_hit|canary_control> (plus ;tokens=<n>;cap=<max> on budget_exceeded); completed;source=<proc_id> — every step of the procedure has run; miss — no stored procedure matched; failed — recall errored and the request was served without it (fail-open). Absent when the skill is off for this key.
anthropic-ratelimit-* / x-ratelimit-* / retry-after / request-id Upstream vendor rate-limit and trace headers relayed verbatim.
MoreThe vendor's own rate-limit headers, relayed verbatim on a response the vendor produced, buffered or streamed. See Vendor rate-limit headers below for the exact list and when they are absent.
x-ace-guardrail / -stages / -mode Injection guard outcome (clean), stages (regex, model), and mode (enforce, shadow).
MoreThe injection guard on a request it did NOT refuse — a refusal is a 400 with x-ace-error: guardrail_violation, so x-ace-guardrail only ever reads clean. -stages says what looked: regex (the pattern firewall) or regex+model (plus the learned classifier). -mode is the CLASSIFIER's stage mode, enforce or shadow, present only when there is a classifier; it is not the skill's mode. x-ace-skill-modes: injection_guard=prod with x-ace-guardrail-mode: shadow is the usual production reading, not a contradiction: the firewall refuses, the classifier scores alongside it and records to telemetry — the deployment is measuring its false-positive rate before letting it refuse anyone. A key running the whole skill in shadow reports shadow here whatever the classifier's own setting. No score in either mode: a per-request P(injection) returned to the sender is a tuning oracle.
x-ace-pii-redacted Count of redacted sensitive entities. In shadow, -would-redact reports estimate.
MoreCount of redacted entities, when PII redaction is on. x-ace-pii-kinds names them (EMAIL=1,PERSON=2); x-ace-pii-stages says which stages ran (ner_regex, ner_model, ner_regex+ner_model). In shadow, x-ace-pii-would-redact carries the count instead and x-ace-pii-redacted is 0.
x-ace-pii-skipped Reports when learned PII stage exceeded its latency budget ceiling (e.g. 50ms).
Morener_model=budget_exceeded: the learned stage ran past its per-request budget (ACE_PII_NER_BUDGET_MS, default 50) and the request was served with the pattern redactions applied and the model's applied only to the turns it finished. ner_model=ceiling_exceeded / overloaded / loop_stalled: the whole scan did not come back inside the skill latency ceiling and the request was served on patterns alone. ner_model=<reason>,ner_regex=<reason> (with x-ace-pii-stages: none): not even the patterns pass answered inside the ceiling and the request went upstream unredacted — named, so a 0 count cannot read as "clean". Every one of these is a verdict on THIS request's wall clock, not a quota: the next request starts from zero. See PII redaction and long prompts and The skill latency ceiling below.
x-ace-compaction-tokens-before / -after / -saved Prompt compaction token metrics (input tokens before, after, and saved).
MorePrompt compaction ran; what went in and what went upstream. x-ace-compaction: refused + -reason when an interlock declined. x-ace-compaction: skipped + -reason (cache_hit, replay_hit) when the skill is on but a stored answer was served above it, so there was no upstream prompt to shrink, or (ceiling_exceeded, overloaded, loop_stalled) when scoring did not finish inside the skill latency ceiling and the prompt went upstream as received — all distinct from absent, which means the skill is off for this key. In shadow, -saved is 0 and -tokens-would-save carries the counterfactual. Under a declared prompt-cache breakpoint the skill runs by the prefix-stable rule and says so with x-ace-compaction-prefix: stable: every rewritable span is pruned by the same pure function of its own bytes on every request (history from a per-process memo, at no compute), so the history goes upstream byte-identical and the first differing byte stays in the live turn. There is no "protect the last N turns" window — a span pruned later than it arrived would move the prefix on every request. The system prompt, which other sessions share, is priced: it is rewritten only where the later cache reads pay for the re-write over the task's horizon; conversation spans are not priced, because ACE wrote their cached bytes on the step they were live (skill_params.prompt_compaction.protect_cached_history: true prices and holds them with the system prompt). Held, x-ace-compaction-prefix: protected with x-ace-compaction-prefix-reason cache_breakpoint_not_worth_it or auto_cache_not_worth_it; priced and pruned anyway, stable with -reason: pays_back. x-ace-compaction-spans: prefix=<n>,user=<n>,tool_results=<n> says where the saving came from. x-ace-compaction-partial: <reason> says part of the prompt went upstream as received although the skill ran, one reason in precedence order: a ceiling shed (ceiling_exceeded / overloaded / loop_stalled — the shed span is pinned as a no-op in the memo so later steps read it at no compute instead of re-scoring it into the same ceiling) outranks ceiling_pinned (a span an earlier step's shed pinned, replayed; the shed itself is reported once) outranks span_too_large (with prune_words on, a span over skill_params.prompt_compaction.max_span_tokens, default 24000) outranks constraints_not_preserved outranks tool_yield_zero (a tool whose last tool_yield_window scored results never pruned is skipped; the spans header then carries a trailing tool_results_skipped=<n>).
x-ace-trajectory-compaction Multi-turn compaction verdict (compacted, would_compact, no_op, refused) and savings.
MoreTrajectory compaction's verdict on this request: compacted, would_compact (shadow), no_op (+ -reason: first_epoch, nothing_sealable, single_turn), refused / skipped (+ -reason: tool_loop, unwritable_body), or failed (fail-open). With it: -mode, -strategy (epoch, the grid fold; sliding_window only for a compactor deployed in that mode), -turns (sealed/total conversational turns), -messages-before / -after, -tokens-before / -after / -saved (shadow: -saved is 0 + -tokens-would-save), and -checkpoint — a 12-hex digest of the summary text that changes exactly when the folded prefix changes, once per epoch, for reconciling against the provider's cache_creation_input_tokens. The fold's state is kept under x-ace-session when sent, else under a handle derived from the transcript's opening. Absent when the skill is off.
x-ace-stage-durations / x-ace-pre-inference-ms / x-ace-full-duration-ms Stage durations, ACE pre-inference latency, and vendor vs gateway wall-clock breakdown.
MoreWhere the wall clock went. x-ace-stage-durations names each stage that completed and, on a 504 gateway_timeout, the one still running — byok_store=30000 is the identity store, upstream=… is the vendor, queue_wait=… is time on the loop before the engine reached the request. Read it FIRST when a request was slow: it separates ACE's own pre-inference work from the provider's, which no other signal does. x-ace-pre-inference-ms is everything ACE did before the upstream call and x-ace-full-duration-ms the whole request, so the difference is the vendor.

Vendor rate-limit headers. When a provider answered — x-ace-served-by names it — its own bucket headers come back with the values it sent, on every vendor-shaped surface: Anthropic's anthropic-ratelimit-{requests,tokens,input-tokens,output-tokens}-{limit,remaining,reset} and request-id; OpenAI's and Azure's x-ratelimit-{limit,remaining,reset}-{requests,tokens}, x-request-id, apim-request-id, x-ms-request-id; retry-after for all. They are the signal a caller spreading load across several upstream keys steers by, and ACE cannot re-derive them. Streamed responses carry them too — they are fixed at connect time, when the upstream status line arrives and before the first frame, the only moment a stream's headers can flush — and so does /anthropic/v1/messages/count_tokens, which draws on the same request bucket. They are absent on any response the vendor did not produce — a semantic-cache hit, an echo turn, an in-fleet completion — because a bucket figure from an earlier call is one a caller would steer by. Framing headers (content-length, content-encoding) are never relayed, and ACE's own 429 rate_limited (the per-tenant cap, no upstream involved) is unchanged. The list is one allowlist shared by every relay path, so it cannot drift between surfaces.

x-ace-warnings codes.

Code Meaning
echo_fallback Served by echo terminal — synthetic reply (0),noupstreamcall.<details><summary>More</summary>Servedbytheechoterminal—asyntheticreply,noupstreamcall,0), no upstream call. <details><summary>More</summary>Served by the echo terminal — a synthetic reply, no upstream call,0. Your key is in echo mode or the request sent x-ace-mode: echo.
model_auto_selected No model requested; auto-selected by gateway router.
MoreNo model was sent and the router chose one; x-ace-route-model says which.
provider_key_pass_through Used request provider key for zero-trust upstream inference.
MoreThe provider key on the request was used for the upstream call (zero-trust), not a stored one.
provider_key_ignored Request key mismatched channel; stored tenant vault key used instead.
MoreA provider key on the request was for a channel other than the one that served, and a stored key served instead. The key you send is meant to be the key that serves you, so a silent substitution is named.
extensions_omitted Provider extension (e.g. thinking block) omitted by terminal or semantic-cache hit.
MoreThe request carried a provider extension the terminal that served it cannot honor, and the REPLY omits it. This covers thinking (a budget, or a replayed signed thinking block) and any other signed or provider-scoped root field. It is emitted from exactly two places — the dispatch path (the echo terminal, a fleet destination; neither can produce a thinking block or a signature) and a semantic-cache hit, where the stored reply predates this request's thinking — and never on a relay to the vendor: a relayed request goes upstream as sent, thinking and signatures included, and comes back as the vendor sent it. So on a live key it appears only on a cache hit; on an echo key it appears on every request carrying thinking, which is the echo terminal saying it could not fake extended thinking, not the gateway saying it stripped yours.
optimization_suppressed All optimization levers forced to shadow by compute_status.
MoreEvery lever was forced to shadow by compute_status.
pipeline_bypassed Internal fault fail-open: request relayed with no levers applied.
MoreA fail-open after an internal fault: the request was relayed with no lever applied.
agent_trajectory_router_reverted Router breaker tripped on validation failures; reverted to fallback tier.
MoreThe per-function escalation breaker tripped: too many cheap-tier answers failed validation, so this function is served from its escalate_to tier until the breaker closes. Also spelled model_router_reverted under the legacy id.

PII redaction and long prompts. pii_ner has two stages. Patterns (email, card, SSN, phone, IP) are linear and always complete. The learned stage — a local NER model for names and account numbers — costs ~28µs a token on the request path, so it is budgeted: ACE_PII_NER_BUDGET_MS (default 50; 0 lifts it) is the most it may spend on one request. Three rules make the budget safe to rely on:

  • A system prompt is not scanned by default, on any surface. It is the operator's instructions, not user data; a name in "you are the assistant for Dr. Okafor's clinic" is the product. Every surface behaves the same way here — Anthropic, Bedrock and Gemini hoist the system prompt, and OpenAI and Responses match them — so the same request measures the same on every door. It is also why a 180k-character system prompt costs the skill nothing at the default. A deployment that does put user data in the system prompt can opt in with ACE_PII_SCAN_SYSTEM (default off). It is deployment-level, like ACE_CACHE_KEY_POLICY: it applies to every tenant on that process and is picked up on redeploy. Switched on, system spans are scanned and redacted on all five surfaces — parity across doors is preserved — and they are scanned BEFORE the conversation, against the same ACE_PII_NER_BUDGET_MS. Only the second property stops holding: a long system prompt now costs the skill what its length implies.
  • The cut lands at the end of the conversation, never in its prefix. Turns are scanned oldest first; when the budget runs out, the turns the model finished keep its redactions and the rest carry the patterns' alone. A conversation grows at the end, so the bytes of the prefix your provider has cached do not change from one turn to the next.
  • What the model has seen, it remembers. Results are memoized per span of text (offsets and labels, never the text; ACE_PII_NER_MEMO_ENTRIES, default 50,000). A replayed history is a lookup, so a 40-turn agent conversation costs the skill a few milliseconds a turn regardless of its length, and a retry of a request the budget cut picks up where the last try stopped rather than starting over.

When the budget does cut a request, the response says so — x-ace-pii-skipped: ner_model=budget_exceeded, x-ace-pii-stages: ner_regex — and the trace stage carries skipped and reason. A deployment whose prompts are routinely cut should raise the budget, pin ACE_PII_NER_THREADS, or narrow ACE_PII_NER_ENTITIES; the number to watch is the skipped header, not the latency.

The skill latency ceiling. Every heavy stage on the request path — pii_ner, semantic_cache (the embed and the vector-store query), prompt_compaction, llm_router and the injection guard's learned stage — runs under a wall-clock ceiling measured by the request itself: 250 ms for a stage the key runs in prod, 50 ms for one it runs in shadow. A stage that has not answered by then is shed: the request is served without it, the response says so on the header that stage already reports through, and the trace stage carries reason, waited_ms and ceiling_ms. It is a bound on one request's wait, not a quota. Nothing accumulates across calls, there is nothing to reset, and the next request is judged on its own clock.

  • ceiling_exceeded — the stage did not answer inside its ceiling; served without it.
  • overloaded — the skill worker queue was already deeper than one ceiling can drain; the stage was not attempted rather than queued to time out later.
  • loop_stalled — the ceiling's timer fired late: the gateway's event loop was held by something else, and the stage may not have run at all. Read this as a gateway fault, not a slow skill — the trace's late_ms is how long the loop did not run.

GET /api/v1/settings reports the effective ceilings under skill_ceiling (per skill, acting and shadow, with the worker pool that runs them); GET /healthz reports the pool as it stands — what is queued, what is running and for how long, and everything shed so far by reason. A gateway shedding at ceiling_exceeded with nothing queued and a job running for minutes has a wedged worker; one shedding at loop_stalled has a held loop; neither is load, and neither is your key.

This header set is frozen legacy: nothing in it is going away without notice, and nothing is being designed onto it — the rows above list what the gateway already sends. A flat string cannot carry what a shadow skill measured, and on a streamed request x-ace-cost-usd is written before the first token exists — an estimate reported as fact. Both are why /v1/execute exists; new skills report on its trace.