all docs
/ reference

What to Send per Skill

What a client sends so each skill can act: the x-ace-session value (your task id) and what reads it, per-skill request inputs (session, use case, task outcome, output and reasoning caps, jurisdiction), what happens without them, and the response header that confirms each skill acted.

Turning a skill on for your key (the settings page, or x-ace-skills per request) decides whether it runs. What your request carries decides whether it has anything to act on. Set these once where your app makes its LLM calls.

x-ace-session

Value. Your own id for one unit of work — a task, an agent run, a conversation — sent unchanged on every call of it: [A-Za-z0-9_-]{1,64}, and a malformed value reads as absent, never as an error. Use the id your application already keeps for the task. Mint a new one for a new task, not for a new step, and keep it when your app rebuilds or summarises its own history mid-task or runs sub-agents in parallel: trajectory compaction keeps its fold state per transcript inside the session, so one session may span several transcripts. Never one id per user or per deployment — two tasks must not share it.

Alternatives. Body session_id is accepted by trajectory compaction and the knowledge graph; on /v1/execute the same handle is attribution.session. Body user is read last and only by trajectory compaction and the knowledge graph. With no session, x-ace-task-id fills the usage row's task id and the router's session scope; trajectory compaction and the knowledge graph do not read it.

Reader What the session does Without it
Usage and request logs The row's session_id and task_id: turns, cost and duration per task in GET /api/v1/usage and the scorecard's task-level numbers. Rows are not grouped by task; the scorecard's task-level fields read null.
agent_trajectory_compaction Scope of the fold state. Still folds: a handle is derived from the transcript's opening messages.
skill_knowledge_graph Groups a task's tool calls for capture; recall tracks its cursor per session. Nothing is captured. Recall still runs, per request.
agent_trajectory_router The session continuity scope: model and effort held for the whole task. No session continuity; see x-ace-archetype.
Canary (canary:N) The bucket: every call of a task gets the same arm. Bucketed per request, so one task mixes arms.

Per skill

Skill Send Without it Confirm on the response
agent_trajectory_compaction A multi-turn body resent in full each step. Recommended: x-ace-session: <task id>. cache_control optional — without one ACE places the prefix marker itself. No session: folded under a derived handle. Fewer than one epoch of turns (5): no_op / first_epoch. x-ace-trajectory-compaction (+ -reason, -turns, -tokens-saved, -write-tokens)
skill_knowledge_graph x-ace-session: <task id> on every call, tool calls in the transcript, and on the task's last call x-ace-task-outcome: success or failure. No session: nothing captured (the derived handle is not used here). No outcome: success is inferred from the absence of errors. A task is extracted after 10 minutes idle. x-ace-knowledge-graph (injected;tokens=N, would_inject, miss, skipped;reason=)
semantic_cache Nothing required. x-ace-use-case per call site lets tool-bearing calls be looked up; x-ace-cache-partition isolates a namespace; Cache-Control: no-cache skips one request. A body with tools or tool history and no use case is not looked up (gated_by_policy); it is still written. Use cases resolving to act or generate are not looked up by default. x-ace-cache (HIT / MISS / BYPASS), -miss-reason, -bypass-reason, -scope
prompt_compaction Nothing. A cache_control breakpoint marks the prefix it must hold byte-stable. — x-ace-compaction (+ -reason, -tokens-saved, -prefix)
llm_router A model ACE knows. x-ace-route-to: <model> pins a call. With a provider key header it swaps within that vendor only. Stands down on signed thinking (continuity:signed) and on any cache_control (continuity:cache_warm). x-ace-route-model, x-ace-route-reason
agent_trajectory_router x-ace-use-case per call site (resolving to a function), a continuity scope (x-ace-session, x-ace-cohort, x-ace-agent-depth or x-ace-archetype), and a routing policy on the key. Keeps your model. A non-integer x-ace-agent-depth also makes it abstain. x-ace-route-*; x-ace-warnings: agent_trajectory_router_reverted
output_budget Your output cap: max_tokens (Anthropic, OpenAI), max_completion_tokens, max_output_tokens (Responses), inferenceConfig.maxTokens (Bedrock), generationConfig.maxOutputTokens (Gemini). no_op / none_declared, unless the key sets a default cap. x-ace-output-budget (+ -requested, -applied, -reason)
reasoning_effort Your reasoning setting: reasoning_effort, reasoning.effort, thinking.budget_tokens (Anthropic; Bedrock under additionalModelRequestFields), thinkingConfig.thinkingBudget. It only ever lowers. none_declared. x-ace-reasoning-effort (+ -requested, -applied, -reason)
feedback_distillation_ring x-ace-use-case per call site, and on accepted answers x-ace-task-status: success or x-ace-user-feedback: positive. Answer and generate pairs without an acceptance signal are skipped (no_acceptance_signal); act calls are never taken. x-ace-distill-ring, x-ace-distill-intake
geo_fence_compliance x-ace-jurisdiction (eu, uk, us, ca, au, unrestricted) or x-ace-data-class (eu-personal, gdpr, uk-personal, phi, hipaa, pci, public). Allowed: the deployment default is normally unrestricted. x-ace-jurisdiction, x-ace-region; x-ace-geo-denied with a 451 on a deny
pii_ner Nothing. Scans turns, tool arguments and tool results; the system prompt only where the deployment enables it. — x-ace-pii-redacted, x-ace-pii-stages, x-ace-pii-kinds
injection_guard Nothing. Scans user turns and tool results, never the system prompt. — x-ace-guardrail (clean, or BLOCKED with a 400)
multi_agent_guard Nothing. Reads tool-call depth and repeats from the transcript. — x-ace-multi-agent-depth; x-ace-multi-agent-halt with a 429
circuit_breaker · adaptive_concurrency · outlier_ejection Nothing. They act on upstream health and load. — x-ace-breaker-open, x-ace-concurrency-limit, x-ace-ejected-upstream

Stack skills (prefix KV cache, LoRA, speculative decoding, quantization and the other GPU-fleet levers) run on infrastructure you operate and need nothing in the request. distillation runs offline.

Header values

Header Send Used for
x-ace-session Your task id, unchanged across the task See above
x-ace-use-case One stable name per call site, e.g. agent_turn, task_memory, classify_state, diagram, or a function name (act, answer, generate, gate, transform, judge, label) Maps to a function: agent_turn→act, classify_state→gate, diagram and task_memory→transform, unless your key's policy says otherwise. Also the key into your fallback chain and the feature in usage
x-ace-task-outcome success or failure, on the task's final call Knowledge-graph capture: a failed task is stored quarantined
x-ace-task-status / x-ace-user-feedback success; positive, accept, thumbs_up Marks an answer as accepted for the distillation ring
x-ace-archetype agentic, qa, generative, pipeline, orchestrator Router continuity and the cache lookup default (qa, pipeline looked up)
x-ace-cohort [A-Za-z0-9_.-]{1,64}, shared by callers of one prefix version Router continuity; a compaction scope when no session is sent
x-ace-agent-depth An integer: the sub-agent's depth Router lineage continuity
x-ace-jurisdiction / x-ace-data-class See geo_fence_compliance Where the request may be served

Minimal agent setup

every call of a task      x-ace-session: <task id>
every call site           x-ace-use-case: <stable call-site name>
the task's final call     x-ace-task-outcome: success | failure
an accepted answer        x-ace-task-status: success

Check x-ace-skill-modes on any response for the mode each skill ran in, then the skill's own header above for what it did.