all docs
/ reference

Provider Surfaces & the Responses API

What every vendor-shaped shim guarantees in both directions — your body goes upstream as written, the vendor's answer (thinking, signatures, reasoning, usage) comes back verbatim, attachments (images, PDFs, audio) reach the provider or are refused by name — plus the OpenAI / Azure Responses API surface and count_tokens.

Every vendor-shaped surface (/v1/chat/completions, /v1/responses, /anthropic/v1/messages, /azure/openai/…, /gemini/v1beta/…, /bedrock/converse and its /bedrock/model/{modelId}/converse alias) is a compatibility shim over the same engine, and each one makes the same two promises.

Request. The body you send is the body the provider receives, byte for byte except inside the text spans a lever rewrote. Thinking configuration and signed thinking blocks, cache_control breakpoints, images, documents, tool definitions, tool results, Gemini thoughtSignature parts and Bedrock reasoningContent all travel as you wrote them. Nothing on this path is parked, projected into another vendor's schema, or silently dropped — and the levers that would invalidate a signature or a cache prefix stand down when one is present.

Response. The provider's answer comes back as the provider sent it, buffered or streamed. On Anthropic that is the full event vocabulary — thinking_delta, signature_delta, redacted_thinking, ping, a mid-stream error, the cache counters on message_start and the vendor's own msg_ id. On Gemini it is every thought part and thoughtSignature, thoughtsTokenCount, cachedContentTokenCount, groundingMetadata and responseId. On Bedrock it is the signed reasoningContent, the cache counters and metrics. Usage is metered from a copy of the stream, so accounting can never alter it. A response the vendor produced also carries the vendor's own rate-limit headers verbatim — anthropic-ratelimit-*, x-ratelimit-*, retry-after and its request id — buffered or streamed; see Response Headers.

Attachments. Images, PDFs and audio reach the provider, whichever provider that is. On a provider surface the part arrives by construction, because the body is relayed as written. When ACE translates — an OpenAI-shaped /v1/chat/completions or /v1/execute request served by Anthropic, Gemini, Vertex or Bedrock — every attachment is rewritten into the destination's schema: image_url becomes an Anthropic image, a Gemini inlineData or a Converse image; a file part (a PDF, base64 or file_url) becomes an Anthropic document, a Gemini inlineData or a Converse document; input_audio becomes a Gemini inlineData. What the destination cannot receive — audio to Anthropic or Bedrock, a URL image to Bedrock (Converse takes bytes, and ACE does not fetch URLs for you), a document type the provider does not read, an OpenAI file_id sent anywhere but OpenAI — is a 400 naming the part and the provider, raised before anything goes upstream. An attachment is never dropped silently. The one place one legitimately leaves a request is trajectory compaction, which replaces whole earlier turns with a summary; the current and retained turns keep theirs.

No cross-provider failover. A request written in one vendor's format is only ever served by that vendor's channel. It never fails forward to another provider, even when your vault holds keys for two channels that serve the same model — an Anthropic Messages body is never posted to Bedrock Converse on an Anthropic 503.

The one exception is POST /v1/messages, the byte-faithful relay: it keeps both promises trivially, by running no lever at all. It exists for parity testing; production Anthropic traffic belongs on /anthropic/v1/messages.

OpenAI Responses API

The OpenAI Responses API is a surface of its own: POST /v1/responses, and for Azure /azure/openai/v1/responses, /azure/openai/responses?api-version=… and /openai/v1/responses (the Azure SDKs have used all three). It is what an OpenAI or Azure client posts for GPT-5 — client.responses.create(...), or the AI SDK's @ai-sdk/openai / @ai-sdk/azure providers, which default to it — and chat completions cannot carry it: the body is item-based, reasoning.summary has no chat equivalent, and store: false + include: ["reasoning.encrypted_content"] is how a Zero-Data-Retention org keeps reasoning continuity across turns. The input[] items (reasoning items with encrypted_content, function_call / function_call_output, input_image and input_file data URLs), tools, tool_choice, reasoning, store, include, max_output_tokens and stream go upstream as sent; the Responses object comes back verbatim, and stream: true relays the provider's own response.* event stream.

A Responses body can only be served by OpenAI or Azure. Pinned to any other provider (x-ace-provider: anthropic, say) it is refused with a 400 naming the reason, in OpenAI's error envelope — never rewritten into chat completions you did not write. A request the semantic cache or an in-fleet model answers is projected into a Responses object (output[] of message and function_call items) so the SDK still parses it; those paths never produce a reasoning item.

from openai import OpenAI

client = OpenAI(
    base_url="https://engine.acefleet.dev/v1",
    api_key="ace_dev_<your-dev-key>",
    default_headers={
        "x-ace-openai-key": "sk-proj-...",  # Zero-Trust header mode
    },
)

# Item-based, not message-based. store=False + include=[...] keeps reasoning
# continuity across turns without OpenAI retaining the conversation.
response = client.responses.create(
    model="gpt-5",
    instructions="You are a concise assistant.",
    input=[{"role": "user", "content": [{"type": "input_text", "text": "Analyze cluster telemetry."}]}],
    reasoning={"effort": "medium", "summary": "auto"},
    store=False,
    include=["reasoning.encrypted_content"],
)
print(response.output_text)

Azure: azure_endpoint pointing at the gateway's /azure prefix, api-key carrying your ACE key (or your Azure key with the ACE key on Authorization: Bearer), the deployment name as model. Upstream, ACE posts to the resource's /openai/v1/responses.

curl -X POST "https://engine.acefleet.dev/azure/openai/v1/responses?api-version=preview" \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-provider: azure" \
  -H "x-ace-azure-key: <your-azure-key>" \
  -H "x-ace-azure-endpoint: https://<resource>.openai.azure.com" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-mini",
    "input": [{ "role": "user", "content": [{ "type": "input_text", "text": "Optimize GPU allocations." }] }],
    "reasoning": { "effort": "low" },
    "store": false,
    "include": ["reasoning.encrypted_content"]
  }'

Anthropic count_tokens

POST /anthropic/v1/messages/count_tokens (and /v1/messages/count_tokens) relays the body to Anthropic with the credential the call carries — the zero-trust header, a non-ACE x-api-key, or your stored Anthropic key — and returns the vendor's count, status included, so a vendor 400 naming a bad block reaches you as such. With no Anthropic credential anywhere the local estimate answers, which cannot count an image, a PDF or a tool schema the way the model's tokenizer does. Echo is decided first: a key in echo mode, or a request sending x-ace-mode: echo, is never relayed whatever credential it carries — the estimate answers with x-ace-served-by: echo and x-ace-warnings: echo_fallback, so a placeholder key sent while testing cannot reach the vendor.

curl -X POST https://engine.acefleet.dev/anthropic/v1/messages/count_tokens \
  -H "x-api-key: ace_dev_<your-dev-key>" \
  -H "x-ace-anthropic-key: sk-ant-..." \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{ "model": "claude-sonnet-4-5", "messages": [{ "role": "user", "content": "ping" }] }'