/ reference
Request Header Specification
Every request header a caller may send: pass-through provider credentials (Bedrock's API key or SigV4 pair included), tenant override, explicit route pinning, the per-request mode (x-ace-mode), the semantic-cache controls (Cache-Control: no-cache, x-ace-cache-partition), the session handle (x-ace-session), the attribution tags (x-ace-team, x-ace-use-case), per-request skill overrides (x-ace-skills) and x-request-id.
Requests authenticate with your ACE developer key. Upstream provider credentials are either passed per-request (zero-trust mode, never persisted) or stored server-side in the provider vault.
| Header | Required | Description | Sample |
|---|---|---|---|
| Authorization | Yes | ACE developer key (ace_dev_...) for tenant auth and telemetry. MoreACE developer key — identity and telemetry attribution for your org. |
Bearer ace_dev_9a8b... |
| x-ace-openai-key | Zero-Trust | Upstream OpenAI API key for pass-through inference (held in memory only). MoreUpstream OpenAI key. Used for this call only, never persisted or logged. |
sk-proj-8f2e... |
| x-ace-anthropic-key | Zero-Trust | Upstream Anthropic API key for pass-through inference (zero-trust). MoreUpstream Anthropic key. Used for this call only, held in memory. |
sk-ant-api03-... |
| x-ace-google-key | Zero-Trust | Upstream Google Gemini API key for pass-through inference. MoreUpstream Google AI key, for the Gemini-shaped surface. |
AIza... |
| x-ace-azure-key | Azure Mode | Upstream Azure OpenAI deployment key. | 7a4b9c1d... |
| x-ace-azure-endpoint | Azure Mode | Azure OpenAI resource endpoint URL. MoreYour Azure OpenAI resource endpoint. |
https:// |
| x-ace-vertex-key | Vertex Mode | GCP service-account JSON (raw or Base64) for Vertex AI. MoreGCP service-account JSON (raw or Base64) for Vertex. The project is auto-extracted from the JSON if not given separately. |
|
| x-ace-vertex-token | Vertex Mode | Direct OAuth bearer token (ya29...) for Google Vertex AI. MoreDirect OAuth bearer token (ya29…) for Google Vertex AI. Alternative to x-ace-vertex-key. |
ya29.a0AfH6SM... |
| x-ace-google-project | Vertex Mode | Google Cloud Project ID for Vertex AI. MoreGoogle Cloud Project ID for Vertex AI (auto-extracted if using Service Account JSON). |
my-gcp-project-123 |
| x-ace-google-region | Optional | Google Cloud region for Vertex AI endpoints (default: us-central1). MoreGoogle Cloud region for Vertex AI regional publisher endpoints (defaults to us-central1).global is addressed at the unprefixed aiplatform.googleapis.com host. |
us-central1 |
| x-ace-bedrock-api-key | Bedrock Mode | Amazon Bedrock API key (ABSK...). Alternative to SigV4; requires region. MoreAn Amazon Bedrock API key (long-termABSK… or short-term), the alternative to the access-key pair below: ACE sends it upstream as Authorization: Bearer and signs nothing, so no AWS secret or STS token leaves your side. Its own header — an ABSK… value on Authorization is an unknown ACE key and a 401. x-ace-aws-region is required with it (a short-term key is bound to the region it was minted in, so ACE refuses to default one). Sent beside x-ace-bedrock-key or x-ace-aws-secret-access-key it is a 400 naming both. Used for this call only, never persisted or logged. |
ABSK... |
| x-ace-bedrock-key | Bedrock Mode | AWS access key ID for SigV4 request signing (computed per request). MoreAWS access key id — half of the SigV4 pair, the alternative to x-ace-bedrock-api-key. ACE computes the SigV4 signature per request, so no local AWS CLI or stored role is needed. |
AKIAIOSFODNN7EXAMPLE |
| x-ace-aws-secret-access-key | Bedrock Mode | AWS secret access key paired with x-ace-bedrock-key. MoreAWS secret access key paired with x-ace-bedrock-key. Never alongside x-ace-bedrock-api-key: the two can belong to different IAM principals, so ACE refuses the combination rather than pick one. |
wJalrXUtnFEMI/K7MDENG/... |
| x-ace-aws-region | Bedrock Mode | AWS region for Bedrock model serving (e.g. us-east-1). MoreAWS region the Bedrock model is served from. Required with x-ace-bedrock-api-key — a missing region is a 400, never a defaulted host; the SigV4 pair alone defaults to us-east-1. |
us-east-1 |
| x-ace-aws-session-token | Optional | Optional AWS session token for temporary STS credentials. MoreAWS session token, for temporary STS credentials. Omit it when signing with a long-lived access key pair. |
IQoJb3JpZ2luX2Vj... |
| x-ace-openai-base-url | Optional | Route OpenAI channel to any OpenAI-compatible host (OpenRouter, Together, vLLM). MoreSend theopenai channel to another OpenAI-compatible host — OpenRouter, Together, Groq, a self-hosted vLLM or SGLang — with that host's key on x-ace-openai-key and its model slug sent verbatim. A trailing /v1 (the SDK convention) is accepted and not doubled. Every channel has the same spelling (x-ace-anthropic-base-url, x-ace-google-base-url, …), and x-ace-base-url is the provider-agnostic form; a stored endpoint on the vendor key is the per-tenant alternative. |
https://openrouter.ai/api/v1 |
| x-ace-base-url | Optional | Provider-agnostic upstream base URL override. MoreProvider-agnostic spelling of the per-channel base-URL override above; the channel's own header wins when both are sent. |
https://openrouter.ai/api/v1 |
| x-ace-vertex-project | Vertex Mode | Alias for x-ace-google-project (Google Cloud project for Vertex AI). MoreAlias of x-ace-google-project. Google Cloud project for Vertex AI, on either surface. |
my-gcp-project-123 |
| x-ace-vertex-region | Optional | Alias for x-ace-google-region (e.g. europe-west1, global). MoreAlias of x-ace-google-region.global is addressed at the unprefixed Vertex host. |
europe-west1 |
| x-ace-provider | Optional | Explicitly pin upstream provider (openai, anthropic, azure, bedrock, google_vertex). MorePin the upstream provider explicitly rather than inferring it from the model. |
azure |
| x-ace-tenant-id | Optional | Explicit tenant namespace override for multi-tenant isolation. | tenant-prod-us |
| x-ace-route-to | Optional | Force a specific provider:model target (e.g. anthropic:claude-sonnet-4-5). MoreForce a specific provider:model pair, overriding engine routing. |
anthropic:claude-sonnet-4-5 |
| x-ace-mode | Optional | Execution mode: live (default) or echo (synthetic 200 for testing, no upstream cost). MoreWhich terminal serves this request.echo (synonyms sandbox, test) answers from the echo terminal — no upstream call, nothing billed; live (synonyms prod, production) forces a real provider call. The per-request value wins over the key's stored echo flag; absent or unrecognised, the key's flag decides, and the default is live. Send live on production traffic so a key left in echo cannot serve a synthetic 200. echo on a deployment without an echo terminal is a 400, never a silent real call. |
echo |
| x-ace-cache-partition | Optional | Isolates semantic cache reads/writes to a named tenant partition. MoreNamespace this request's semantic-cache reads and writes to a partition of your org's pool that nothing else shares ([A-Za-z0-9_-]{1,64}). It only ever narrows — a partition cannot reach the org-wide pool, another partition or another tenant — so benchmark and load-test traffic can run against a production tenant without writing into the pool that serves real users. A present-but-malformed value fails closed to a namespace unique to the request (always MISS) rather than falling through to the shared pool. The response confirms the scope with x-ace-cache-scope. |
loadtest-2026-09 |
| x-ace-session | Optional | Your ID for one task or agent run: groups its calls for reporting and compaction. MoreAn opaque handle minted by your side ([A-Za-z0-9_-]{1,64}; a malformed value reads as absent) naming one unit of work — a task, an agent run, a conversation — sent unchanged on every call of it. Use the id your application already has for that unit (a task id is the usual choice). It does not have to change when your app rebuilds or summarises its own history: agent-trajectory compaction keeps its fold state per transcript inside the session, so one session may span several transcripts, in turn or side by side. On /v1/execute the same handle is attribution.session, and agent-trajectory compaction also accepts a body session_id. Three things read it: telemetry, where it is the session_id on the request row; agent-trajectory compaction, which keys its fold state on it — without one, the gateway derives a handle from the transcript's opening messages, which every call of one transcript re-sends, so sending none still folds; and, on a deployment running the semantic cache in session-isolation mode (public demos, not the production default), the cache namespace — there a request without one is isolated to itself and always misses. It is a reporting label: it never selects the tenant, the key or a rate-limit bucket. |
sess_7f3a9c2e |
| x-ace-team | Optional | Team attribution tag for spend reporting and per-team budget limits. MoreAttribution tag: theteam label on this request's telemetry row, and the team dimension of GET /api/v1/usage. Absent, the row carries your key's public stable id instead, so untagged traffic still groups by key. On /v1/execute the same label is attribution.tenant. It keys the per-team budget window, and nothing else: it does not move the semantic-cache namespace, the virtual-key id or the rate-limit bucket, and a label naming another tenant on a dedicated deployment is a 403, not a relabel. |
payments-platform |
| x-ace-use-case | Optional | Feature tag (e.g. support-bot) for fallback chains and usage grouping. MoreAttribution tag: thefeature dimension of GET /api/v1/usage (use_case on the row; attribution.use_case on /v1/execute). The one attribution field with a serving consequence — it is also the key into your tenant's stored fallback chain, so a tagged request has a fallback where an untagged twin has none. Send one per product surface (support-bot, code-review) and the usage rollup separates their spend without a second key per feature. |
support-bot |
| x-ace-archetype | Optional | Agent routing preset: agentic, qa, generative, pipeline, orchestrator. MoreWhich routing preset the agent trajectory router applies to this request:agentic, qa, generative, pipeline or orchestrator (attribution.archetype on /v1/execute). Sets the continuity scope and the cache strategy the function policy starts from; absent, the key's configured default applies, and with none the request is continuity: none and is never purpose-routed. |
agentic |
| x-ace-cohort | Optional | Prefix version hash or knowledge-base build for sticky canary routing. MoreA prefix-version id shared by many callers — a knowledge-base build, a system-prompt hash (attribution.cohort on /v1/execute). The agent trajectory router pins model and effort per cohort and moves them only when the cohort's version changes or through a sticky canary bucket, so a shared cache is never re-warmed for everyone at once. |
kb_v2026_09 |
| x-ace-skills | Optional | Per-request skill overrides (skill_id=mode, e.g. *=off,injection_guard=prod). MorePer-request skill overrides on every vendor-shaped surface (and, beneath its ownskill_overrides field, on /v1/execute): comma-separated skill_id=mode pairs, mode one of off / shadow / prod — the same catalogue, vocabulary and validator as the native override, strict and case-sensitive. canary:N and rolled_back are not per-request modes: a canary is a fraction of a key's traffic, so it is set on the key through the lifecycle route (see the endpoint reference). The override beats your stored mode in either direction; nothing is persisted. An unknown skill id or mode is a 400 in the surface's own envelope naming the pair, refused before dispatch; an override on a skill your org has locked is a 403 skill_locked projected per surface (Anthropic permission_error, Gemini PERMISSION_DENIED, Bedrock AccessDeniedException) with the locked ids in the message. On the OpenAI shim a skill named in extra_body.ace.skill_overrides beats the header's pair for that skill. The response's x-ace-skills-applied echoes the pairs that actually ran. The valid ids are exactly the universal list GET /api/v1/skills returns (every skill whose scope is SKILL_SCOPE_UNIVERSAL and is not deprecated — the same set POST /api/v1/dev_key/skills accepts); it includes skill_knowledge_graph, output_budget, reasoning_effort and agent_trajectory_router (legacy alias model_router), and it grows as skills ship, so a hand-maintained off-list goes stale: a control arm that names the seven skills it knew about leaves the eighth running and measures a system with one lever still pulled. Use the wildcard instead: *=off (or *=shadow / *=prod) stands for every universal skill the gateway registers, resolved per request from its registry, never from a list; the same "*": {"mode": "off"} entry is accepted in skill_overrides. A pair naming a skill explicitly in the same header or body beats the wildcard for that skill (*=off,injection_guard=prod), whichever side each came from. * with no mode, or a mode outside the vocabulary, is the same 400. A skill your org has locked is skipped by the wildcard rather than refused — you did not name it — and reported as <skill>=locked in x-ace-skills-applied; naming a locked skill beside * is still the 403. There is no all= alias. |
*=off,injection_guard=prod |
| Cache-Control | Optional | no-cache forces cache bypass and direct upstream provider execution. Moreno-cache (exactly that value): this request is never served from the semantic cache. The provider still answers and the answer is still stored for later hits; only the read is skipped. Use it on a call whose answer must be fresh — a grounded search, a tool-bearing turn. A request with temperature > 0.8 is not served from the cache either. On /v1/execute, execution.no_cache is the same switch, and an explicit false there overrides a stale header a client library set once. |
no-cache |
| x-request-id | Optional | Custom client request ID, propagated to logs and responses. MoreYour own id for this request. It becomes the gateway's request id — theid of a /v1/execute response, the request_id in the request log, and the tag on every log line for the call — in place of a generated req-…. It is not echoed as a response header on the vendor-shaped surfaces. |
order-7f3a-retry-2 |