all docs
/ reference

Endpoint Reference

The customer-facing endpoint surface: the native /v1/execute engine, the OpenAI chat-completions and Responses API shims, the Anthropic (messages and count_tokens), Azure, Gemini and Bedrock shims (including the /bedrock/model/{modelId}/converse paths an AWS SDK builds), self-service account configuration, per-key skill modes and the lifecycle route (off/shadow/canary:N/prod/rolled_back), the skill scorecard and its breaker thresholds, and the tenant-scoped reads: usage by dimension (/api/v1/usage), the request inspector feed and per-request logs.

Base URL: https://engine.acefleet.dev

Method Path Auth Description
POST /v1/execute Dev key The native endpoint. Structured input/target/execution envelope, typed SSE streaming, per-request skill overrides, and a trace that reports what every skill did — including skills running in shadow, which no response header can express.
POST /v1/chat/completions Dev key OpenAI-compatible shim. Drop-in for the OpenAI SDK by changing base_url.
POST /v1/responses Dev key OpenAI Responses API shim — what the OpenAI SDK's client.responses.create and the AI SDK's @ai-sdk/openai post for GPT-5. The item array, reasoning, store, include and encrypted_content go upstream as sent; the Responses object and its response.* stream come back verbatim. Served by OpenAI or Azure only.
POST /azure/openai/v1/responses Dev key Azure Responses API shim. Accepts the ACE dev key on api-key, the deployment name as model, and the api-version query the Azure SDKs send. /azure/openai/responses?api-version=… and /openai/v1/responses are the same route.
POST /anthropic/v1/messages Dev key Anthropic-compatible shim. Accepts the ACE dev key on x-api-key, so the Anthropic SDK needs no code around it. Thinking, tools, images and cache_control are relayed as sent; the reply — and the stream, thinking blocks and signatures included — is Anthropic's own. With x-ace-provider: google_vertex and Vertex credentials the same call serves Claude on Vertex (:rawPredict) — see the Google Vertex AI provider integration.
POST /anthropic/v1/messages/count_tokens Dev key Token count for a Messages request, relayed to Anthropic with the call's credential and returned as the vendor answered. Falls back to a local estimate only when no Anthropic credential is available.
POST /v1/messages Dev key Not for production use — send Anthropic traffic to /anthropic/v1/messages instead. A byte-faithful relay built for parity testing: the body is forwarded as the exact bytes received, so every ACE lever is off (no semantic cache, prompt compaction, PII redaction, model routing or merit-order dispatch). It preserves however you presented your credential and never rewrites it.
POST /azure/openai/… Dev key Azure OpenAI-compatible shim. Accepts the ACE dev key on api-key; the AAD/Entra bearer path stays intact.
POST /gemini/v1beta/models/{model}:generateContent Dev key Google Gemini Developer API and Vertex AI native shim — the same GenerateContentRequest Google's own endpoints take, relayed as written. Served by the Gemini Developer API unless pinned with x-ace-provider: google_vertex, which sends it to Vertex's publishers/google under the project and region on the request. :streamGenerateContent is the streaming counterpart. thought parts and thoughtSignature travel both ways as sent, so a signed function call replays across agent turns.
POST /bedrock/converse Dev key AWS Bedrock Converse proxy for a hand-written caller, modelId in the body. Sends your Converse body — document, image, cachePoint, reasoningContent and all — upstream with an Amazon Bedrock API key as a bearer, or SigV4-signed with the access-key pair, and returns the vendor's reply. /bedrock/converse-stream is the streaming counterpart.
POST /bedrock/model/{modelId}/converse Dev key The same surface at the path an AWS SDK or the AI SDK's Bedrock provider builds from a base URL of /bedrock: modelId is read from the path, percent-encoded as the SDK sends it (: as %3A; an inference-profile or custom-model ARN with / as %2F), and lifted into the body when the body has none. A body that repeats it proceeds; a body that names a different id is a 400 ValidationException naming both. /bedrock/model/{modelId}/converse-stream is the streaming counterpart.
GET /v1/models None List models the fleet can currently serve. No auth.
GET /healthz None Liveness, plus the active lever set for this deployment — the fastest way to see what is actually switched on.
POST /api/v1/vendor_key/create Dev key Store an upstream provider key server-side, so you can stop sending pass-through headers. For bedrock, api_key is the packed access-key pair by default; credential_kind: "api_key" stores an Amazon Bedrock API key instead.
POST /api/v1/vendor_key/delete Dev key Remove a stored provider key.
GET /api/v1/settings Dev key Read your org's current configuration.
POST /api/v1/spend_cap/update Dev key Set a real-time spend cap.
POST /api/v1/model_access/update Dev key Change which models are enabled for your org.
POST /api/v1/dev_key/skills Dev key Set one developer key's skills. Body {dev_key_id, skills: {semantic_cache: true, …}, skill_modes: {semantic_cache: "shadow", …}, locked_skills?: {…}, skill_params?: {…}}; dev_key_id is the console's key id, or the gateway's key_id for a key minted on-premises. Flags and modes are merged, so one toggle does not clear the rest, and a skill never set is off — no deployment setting turns a skill on for a key. locked_skills refuses per-request x-ace-skills overrides for that skill with a 403 skill_locked; a lock must accompany or follow a stored mode. skill_params (per-skill numeric policy, today agent_trajectory_compaction) is Enterprise-only: on any other plan the whole save is a 403 naming it. The next request served picks the write up.
POST /api/v1/tenant/{org_id}/skills/{skill_id}/lifecycle Dev key Move a skill through its lifecycle for one key. Body {key_id, mode, reason?, trigger_source?}; mode is one of off / shadow / canary:N (N in 1–99; bare canary means 10) / prod / rolled_back, anything else a 400 naming the vocabulary. key_id is required: the gateway's key_id or the console's dev_key_id, for a key in org_id or one of its sub-tenants (404 otherwise); skill_id must be a per-key skill (400 otherwise); trigger_source defaults to user:api. Writes the key's stored mode — the same store /api/v1/dev_key/skills writes — and the next request served uses it; a write that fails is an error, never a 200. Returns {status: "success", stored_under, skill_modes, transition: {id, org_id, key_id, skill_id, from_mode, to_mode, canary_percent, trigger_source, reason, timestamp}} — skill_modes is the key's stored modes after the write. See Skill lifecycle, canary and auto-revert below.
GET /api/v1/tenant/{org_id}/skills/{skill_id}/scorecard Dev key The production scorecard for one skill: ?window= (15m, 1h — the default — 24h, 7d) and ?key_id= to narrow to one key. Computed over the durable rows when there is a store (source.store: durable, every replica's traffic), over the answering process's rolling window only without one. Returns status (HEALTHY, REGRESSION_WARNING, BREAKER_ADVISORY — a tripwire crossed that this deployment will not act on — CIRCUIT_BREAKER_TRIGGERED — crossed and the key is being rolled back — or NO_DATA when nothing ran; an unmeasured skill is never reported healthy), breaker_action (auto_rollback | advisory), breakers[] (the machine names of the tripwires that crossed), confounded + confounded_reason (see the lifecycle note), sample_size {total, treatment, control}, generic_metrics (gateway overhead and full-duration percentiles, upstream error, client retry and fail-open rates, treatment and control side by side), downstream_ai_metrics (turns per task, e2e duration, cost per task, task success and goodput — computed over the samples that carry a session/task id only, per arm, with session_share_*, task_sample_size and a single_call_metrics block for the rest; tool-call fidelity and client friction per request), skill_specific_metrics, anomalies[] naming each floor crossed, and deployment_history — the last 20 lifecycle transitions.
GET /api/v1/usage Dev key Usage rolled up by any whitelisted dimension over the durable request rows: ?group_by=provider,team&window=7d&bucket=day&limit=100&key_id=…. group_by (comma-separated, default provider) accepts endpoint, provider, channel, model, team (x-ace-team), feature (x-ace-use-case), key, status, cache_tier, tier, egress_mode, relay_reason, optimization_state, route, desk, pm, strategy, cost_centre (or cost_center) and project; bucket is hour, day or week; any other dimension is a drill-down filter (?provider=azure). Rows carry requests, cost_usd, tokens_in, tokens_out and cache_hits. Your org is implied by the credential — org is neither a dimension nor a filter (400) — and key_id is verified before it filters. Needs a telemetry store; a deployment without one answers 501 rather than an empty table.
GET /api/v1/requests Dev key The per-request inspector feed: the last 100 requests for exactly one of ?key_id= or ?tenant_id= (400 with neither or both; limit ≤ 100), filterable on model, endpoint, provider, query_category, feature and cache. Each row is the cache result, the routed model and reason, tokens in/out/saved and the cost, plus a prompt preview — which is why it is authenticated. An in-memory ring per process: nothing here persists past 100 requests, and on a deployment with more than one replica it is the answering replica's slice of your traffic — a tenant whose calls landed elsewhere sees an empty list. The durable list is GET /api/v1/tenant/requests.
GET /api/v1/tenant/requests Dev key The durable inspector list: your tenant's request_log rows in ?window= (default 1h), newest first, ?limit= ≤ 500, from the store every replica writes to — so it is the list /api/v1/requests cannot be, and the discovery path for a request id to hand to /api/v1/tenant/request/{request_id}. Same allowlisted shape as that read (no bodies); ?decision=true carries each row's decision record inline. Filters are any allowlisted column as a query parameter, matched exactly: ?session_id= lists one run's turns, ?key_id=, ?model=, ?provider=, ?surface=, ?channel=, ?tier=, ?relay_reason=, ?optimization_state=, ?status=, ?cache_tier=, ?team=, ?use_case=, ?destination_id=, ?model_serving_route=, ?egress_mode=; an unknown filter is a 400 naming the allowed columns. The tenant is resolved from the credential and is the row fetch's own predicate, so no filter widens the result past your rows. source names the store that answered. Needs a telemetry store (501 without one).
GET /api/v1/tenant/request/{request_id} Dev key One request's durable row and its decision record — what the gateway decided and why, after the fact, without the bodies. Tenant-scoped from the credential; the id is the gateway's x-ace-request-id (your x-request-id when you sent one), found via /api/v1/tenant/requests.
GET /api/v1/logs/request/{request_id} Dev key One request's gateway log lines, ?tenant_id= required (after, limit ≤ 2000, level, q, logger narrow further). The id must be in your tenant's own last-100 inspector ring: a request you do not own and one that never existed are both 404, so the endpoint is not an oracle for other tenants' ids, and a request older than the ring has aged out. request_id is the gateway's id — your x-request-id when you sent one.

Self-service configuration endpoints resolve identity from your own dev key, scoped to your own org. An org key naming a different tenant_id in the body is rejected with a 403, never silently overridden.

Who may read a tenant's telemetry. /api/v1/usage, /api/v1/requests, /api/v1/logs/request/{id} and the tenant scorecard and lifecycle routes are authorized by one rule: the credential must be entitled to configure the tenant it names. That is your org's own dev key (an org key is scoped to its own tenant, and naming another is a 403), or a deployment's config secret or admin key naming the tenant explicitly. No key at all is a 401. key_id is a public identifier, not a credential — it is resolved to the tenant that owns it and authorized the same way, so an unknown id is refused rather than waved through. Authorization is checked before existence, which is what keeps a 404 from confirming that someone else's id is real.

Skill lifecycle, canary and auto-revert. A skill has five stored modes per key: off, shadow (the decision runs on real traffic and is thrown away — measured, never acted on), canary:N (the rollout mode: the skill acts, and the scorecard reports its treatment branch against a control branch), prod and rolled_back (a quarantine state a breaker or an operator puts a skill in; nothing promotes out of it automatically). The lifecycle route and /api/v1/dev_key/skills write the same per-key store, so a transition posted here is what the next request runs under.

Per request, only off, shadow and prod. The x-ace-skills header (and skill_overrides on /v1/execute) substitutes a mode into one request; that is what it is for — a CI job proving a skill's effect, one request run in shadow. canary:N is not in that vocabulary because it is a statement about a population: a fraction of a key's requests, with the rest serving as the control the scorecard measures against. A single request cannot be 10% of itself, and rolled_back is a verdict about a rollout, not an instruction for a call. Both belong to the key, so both are set on the lifecycle route.

What trips the scorecard breaker. The scorecard's status is computed over the window on every read, once at least 20 samples exist (and, for any treatment-vs-control comparison, 20 on each branch). A tripwire crosses when: fail-open rate is above 1.0% of requests (an optimization that timed out or stalled and was served unoptimized); tool-call schema fidelity in treatment is below 97%; downstream turns per task are up more than 25% in treatment over control; task success rate is more than 5 points below control; goodput ratio is more than 5 points below control; or, for agent_trajectory_compaction, the fold's net dollar ledger is negative over 5 or more fold moves. Each one is named in breakers[] and explained in anomalies[]. It is REGRESSION_WARNING — reported, not tripped — when the upstream error rate in treatment exceeds control by more than 5 percentage points, when the client retry rate in treatment is more than 2× control, or when turns per task are up 15–25%.

The task-level tripwires compare like with like. Turns per task, e2e duration, cost per task, task success and goodput are computed only over requests that carry a session or task id (x-ace-session), in both arms; a single call is not a one-turn task, and an arm with no sessions reads null for those fields rather than a perfect score. When the two arms' session coverage differs by more than 20 percentage points the summary reports confounded: true with a confounded_reason such as session_share_treatment=1.00 control=0.00, the task-level numbers are still shown, and the task-level tripwires do not fire — a suite that sends x-ace-session only from its multi-step workflows, all in prod, would otherwise read "turns per task +347%" under every skill at once. Send x-ace-session on both arms of a comparison.

CIRCUIT_BREAKER_TRIGGERED means the breaker will act. The status a crossed tripwire renders depends on whether this deployment's auto-rollback is armed for the skill (ACE_SKILL_AUTO_ROLLBACK; by default the trajectory fold alone): armed, it is CIRCUIT_BREAKER_TRIGGERED and breaker_action: auto_rollback, and the gateway's once-a-minute background sweep — never a request — moves the key's stored mode to rolled_back and records the transition with trigger_source: auto:<breaker> (e.g. auto:fail_open_rate, joined with + when several crossed) and the anomalies as its reason. Not armed, the same numbers read BREAKER_ADVISORY with breaker_action: advisory: nothing changes on the key, and the scorecard never claims it did. Either way, watch deployment_history for what actually happened; posting rolled_back to the lifecycle route yourself is the same transition under your own trigger_source.