all docs
/ reference

Integration Quickstarts

Copy-pasteable OpenAI (chat completions and the Responses API), Anthropic, Azure, Gemini, Vertex, Bedrock and native-engine integrations, and the Vercel AI SDK providers (including Claude on Vertex through the Anthropic provider) — the drop-in is a base URL change plus one header.

ACE is a drop-in reverse proxy. Point a standard SDK at the gateway with your ACE developer key and pass the upstream provider credential as a header.

OpenAI SDK (Python)

import openai

client = openai.OpenAI(
    base_url="https://engine.acefleet.dev/v1",
    api_key="ace_dev_<your-dev-key>",
    default_headers={
        "x-ace-openai-key": "sk-proj-...",  # Zero-Trust header mode
    },
)

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Analyze cluster telemetry."}],
)
print(response.choices[0].message.content)

OpenAI proxy (cURL)

curl -X POST https://engine.acefleet.dev/v1/chat/completions \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-openai-key: sk-proj-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{ "role": "user", "content": "Analyze cluster telemetry." }]
  }'

OpenAI Responses API (Python)

from openai import OpenAI

client = OpenAI(
    base_url="https://engine.acefleet.dev/v1",
    api_key="ace_dev_<your-dev-key>",
    default_headers={
        "x-ace-openai-key": "sk-proj-...",  # Zero-Trust header mode
    },
)

# Item-based, not message-based. store=False + include=[...] keeps reasoning
# continuity across turns without OpenAI retaining the conversation.
response = client.responses.create(
    model="gpt-5",
    instructions="You are a concise assistant.",
    input=[{"role": "user", "content": [{"type": "input_text", "text": "Analyze cluster telemetry."}]}],
    reasoning={"effort": "medium", "summary": "auto"},
    store=False,
    include=["reasoning.encrypted_content"],
)
print(response.output_text)

Azure Responses API (cURL)

curl -X POST "https://engine.acefleet.dev/azure/openai/v1/responses?api-version=preview" \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-provider: azure" \
  -H "x-ace-azure-key: <your-azure-key>" \
  -H "x-ace-azure-endpoint: https://<resource>.openai.azure.com" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-mini",
    "input": [{ "role": "user", "content": [{ "type": "input_text", "text": "Optimize GPU allocations." }] }],
    "reasoning": { "effort": "low" },
    "store": false,
    "include": ["reasoning.encrypted_content"]
  }'

Anthropic proxy (cURL)

curl -X POST https://engine.acefleet.dev/anthropic/v1/messages \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-anthropic-key: sk-ant-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [{ "role": "user", "content": "Optimize GPU allocations." }]
  }'

Azure OpenAI proxy (cURL)

curl -X POST "https://engine.acefleet.dev/azure/openai/deployments/gpt-4o-prod/chat/completions?api-version=2024-02-15-preview" \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-provider: azure" \
  -H "x-ace-azure-key: <your-azure-key>" \
  -H "x-ace-azure-endpoint: https://<resource>.openai.azure.com" \
  -H "Content-Type: application/json" \
  -d '{ "messages": [{ "role": "user", "content": "Optimize GPU allocations." }] }'

Google Gemini proxy (cURL)

curl -X POST "https://engine.acefleet.dev/gemini/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "x-goog-api-key: ace_dev_<your-dev-key>" \
  -H "x-ace-provider: google_gemini" \
  -H "x-ace-google-key: <your-gemini-key>" \
  -H "Content-Type: application/json" \
  -d '{ "contents": [{ "role": "user", "parts": [{ "text": "Optimize GPU allocations." }] }] }'

Google Vertex AI proxy (cURL)

curl -X POST "https://engine.acefleet.dev/gemini/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "x-goog-api-key: ace_dev_<your-dev-key>" \
  -H "x-ace-provider: google_vertex" \
  -H "x-ace-vertex-token: ya29..." \
  -H "x-ace-google-project: <your-gcp-project>" \
  -H "x-ace-google-region: us-central1" \
  -H "Content-Type: application/json" \
  -d '{ "contents": [{ "role": "user", "parts": [{ "text": "Optimize GPU allocations." }] }] }'

AWS Bedrock proxy (cURL)

# An Amazon Bedrock API key: sent upstream as a bearer, nothing signed, no AWS secret
# leaves your side. The region is required with it.
curl -X POST https://engine.acefleet.dev/bedrock/converse \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-bedrock-api-key: ABSK..." \
  -H "x-ace-aws-region: us-east-1" \
  -H "Content-Type: application/json" \
  -d '{
    "modelId": "anthropic.claude-3-5-sonnet-20241022-v2:0",
    "messages": [
      { "role": "user", "content": [{ "text": "Optimize GPU allocations." }] }
    ]
  }'

# The SigV4 pair instead (never both): ACE signs the request for the AWS host.
# modelId may also travel in the path, as an AWS SDK sends it (percent-encoded).
curl -X POST https://engine.acefleet.dev/bedrock/model/anthropic.claude-3-5-sonnet-20241022-v2%3A0/converse \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-bedrock-key: AKIAIOSFODNN7EXAMPLE" \
  -H "x-ace-aws-secret-access-key: wJalrXUtnFEMI/K7MDENG/..." \
  -H "x-ace-aws-region: us-east-1" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      { "role": "user", "content": [{ "text": "Optimize GPU allocations." }] }
    ]
  }'

ACE native engine (cURL)

curl -X POST https://engine.acefleet.dev/v1/execute \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "Content-Type: application/json" \
  -d '{
    "input":  { "messages": [{ "role": "user", "content": "Audit node health." }] },
    "target": { "model": "gpt-4o" },
    "execution": { "trace": "full" }
  }'

Vercel AI SDK providers

An AI SDK provider is configured through the three knobs every createX() factory exposes — baseURL, headers and fetch — and on ACE the integration is the first two. baseURL moves to the gateway's vendor-shaped surface for that provider; the ACE developer key goes in apiKey, because each surface reads it from the header the provider already sends (x-api-key, api-key, x-goog-api-key or Authorization: Bearer); the upstream credential rides on the zero-trust header in headers. fetch is not needed on any provider: every path the six factories build, Bedrock's /model/{modelId}/converse[-stream] included, is one the gateway serves. The body the SDK serializes is what the vendor receives and the vendor's reply is what the SDK parses, streamed or not — the guarantees in Provider Surfaces apply unchanged.

Anthropic (@ai-sdk/anthropic). apiKey is sent as x-api-key. The shim treats an ace_dev_ value there as your ACE identity and fills Authorization: Bearer with it when you set none; any other value is a pass-through Anthropic key. So the dev key goes in apiKey and the vendor key on x-ace-anthropic-key — or, equivalently, the vendor key in apiKey and Authorization: Bearer ace_dev_… in headers. Thinking, tools, images and cache_control are relayed as sent; the reply and its stream are Anthropic's own. Token counting is a plain POST of the same body and headers to /anthropic/v1/messages/count_tokens, which returns the vendor's count.

import { createAnthropic } from '@ai-sdk/anthropic';

const anthropic = createAnthropic({
  baseURL: 'https://engine.acefleet.dev/anthropic/v1',
  apiKey: 'ace_dev_<your-dev-key>',                 // sent as x-api-key: your ACE identity
  headers: { 'x-ace-anthropic-key': 'sk-ant-...' }, // zero-trust: used for this call only
});
// model: anthropic('claude-sonnet-4-5')

OpenAI and Azure (@ai-sdk/openai, @ai-sdk/azure). Both default to the Responses API, which the gateway serves as its own surface — /v1/responses, and under /azure/openai the three spellings the Azure SDKs use — with the item array, reasoning, store, include and encrypted_content relayed as sent and the response.* stream returned verbatim; chat completions from the same factories land on /v1/chat/completions and /azure/openai/deployments/{deployment}/chat/completions. @ai-sdk/openai sends apiKey as Authorization: Bearer; @ai-sdk/azure sends it as api-key, which the Azure surface reads the same way. Azure has no global host, so the resource endpoint travels on x-ace-azure-endpoint and the deployment name is the model.

import { createOpenAI } from '@ai-sdk/openai';
import { createAzure } from '@ai-sdk/azure';

const openai = createOpenAI({
  baseURL: 'https://engine.acefleet.dev/v1',
  apiKey: 'ace_dev_<your-dev-key>',                 // Authorization: Bearer
  headers: { 'x-ace-openai-key': 'sk-proj-...' },
});

const azure = createAzure({
  baseURL: 'https://engine.acefleet.dev/azure/openai',
  apiKey: 'ace_dev_<your-dev-key>',                 // sent as api-key
  headers: {
    'x-ace-provider': 'azure',
    'x-ace-azure-key': '<your-azure-key>',
    'x-ace-azure-endpoint': 'https://<resource>.openai.azure.com',
  },
});
// model: azure('gpt-4o-prod') — the deployment name

OpenRouter and other OpenAI-compatible hosts. Any OpenAI-compatible host is the openai channel pointed elsewhere: x-ace-openai-base-url names the host, x-ace-openai-key carries its key, and the model slug is sent verbatim — anthropic/claude-sonnet-4.5 reaches OpenRouter spelled exactly like that. x-ace-served-by still names the channel (openai), not the host. The same headers work from @openrouter/ai-sdk-provider or any provider package whose factory takes baseURL and headers, pointed at /v1. A Responses body on this channel is posted to {host}/responses, so a host that speaks only chat completions must be sent chat completions.

import { createOpenAI } from '@ai-sdk/openai';

const openrouter = createOpenAI({
  baseURL: 'https://engine.acefleet.dev/v1',
  apiKey: 'ace_dev_<your-dev-key>',
  headers: {
    'x-ace-provider': 'openai',
    'x-ace-openai-base-url': 'https://openrouter.ai/api/v1',
    'x-ace-openai-key': 'sk-or-v1-...',
  },
});
// model: openrouter.chat('anthropic/claude-sonnet-4.5') — the host's own slug

Claude on Vertex, through the Anthropic shim. The one callers miss: no Vertex-specific package is needed. Claude on Vertex takes an Anthropic Messages body, which the Anthropic provider already produces, so the channel is only a different address and credential. Pinned with x-ace-provider: google_vertex, the same call goes to …/publishers/anthropic/models/{model}:rawPredict (:streamRawPredict when streaming): model moves out of the body into the path — use Vertex's id, e.g. claude-sonnet-4-5@20250929 — anthropic_version: vertex-2023-10-16 is set in the body, an anthropic-beta header travels as the anthropic_beta array, and the reply, thinking and signatures included, is Anthropic's verbatim. One of x-ace-vertex-token (a ya29 bearer) or x-ace-vertex-key (service-account JSON) is required, as is x-ace-google-project; x-ace-google-region defaults to us-central1.

import { createAnthropic } from '@ai-sdk/anthropic';

const claudeOnVertex = createAnthropic({
  baseURL: 'https://engine.acefleet.dev/anthropic/v1',
  apiKey: 'ace_dev_<your-dev-key>',
  headers: {
    'x-ace-provider': 'google_vertex',
    'x-ace-vertex-token': 'ya29...',          // or 'x-ace-vertex-key': '<service-account JSON>'
    'x-ace-google-project': '<your-gcp-project>',
    'x-ace-google-region': 'europe-west1',
  },
});
// model: claudeOnVertex('claude-sonnet-4-5@20250929') — Vertex's id, sent in the path upstream

Gemini on Vertex, through the Gemini shim. /gemini/v1beta/models/{model}:generateContent (and :streamGenerateContent) takes Google's own GenerateContentRequest, served by the Gemini Developer API or — with x-ace-provider: google_vertex and the Vertex credential set above — by Vertex's publishers/google surface; the body goes upstream as written and the reply, thought parts and thoughtSignature included, comes back as sent. The credential contract is exact: the ACE dev key is read from Authorization: Bearer, or from x-goog-api-key when no bearer is present — a bearer you set yourself is never overwritten. A client that puts a Google OAuth token on Authorization: Bearer, the shape a Vertex-native client produces, therefore leaves the request with no ACE identity and is refused; whether a createVertex-built request can be made to fit is not something this reference claims. The Vertex token belongs on x-ace-vertex-token, and the dev key on x-goog-api-key (as @ai-sdk/google's createGoogleGenerativeAI sends apiKey) or on the bearer.

import { createGoogleGenerativeAI } from '@ai-sdk/google';

const geminiOnVertex = createGoogleGenerativeAI({
  baseURL: 'https://engine.acefleet.dev/gemini/v1beta',
  apiKey: 'ace_dev_<your-dev-key>',           // sent as x-goog-api-key: your ACE identity
  headers: {
    'x-ace-provider': 'google_vertex',
    'x-ace-vertex-token': 'ya29...',          // or 'x-ace-vertex-key': '<service-account JSON>'
    'x-ace-google-project': '<your-gcp-project>',
    'x-ace-google-region': 'us-central1',
  },
});
// Without x-ace-provider and the Vertex headers, the same client is served by the
// Gemini Developer API with 'x-ace-google-key': '<your-gemini-key>'.

Bedrock (@ai-sdk/amazon-bedrock). baseURL is the gateway's /bedrock prefix and nothing else changes: the provider builds {baseURL}/model/{modelId}/converse[-stream] with the id percent-encoded in the path (: as %3A; an inference-profile ARN with / as %2F), and the gateway serves exactly that path, lifting modelId from it into the body. No request-rewriting fetch. The upstream credential is one of two on the zero-trust headers, never both: an Amazon Bedrock API key on x-ace-bedrock-api-key, which ACE sends upstream as a bearer with nothing signed, or the access-key pair on x-ace-bedrock-key + x-ace-aws-secret-access-key, which ACE SigV4-signs for the real AWS host. x-ace-aws-region names the region either way (required with the key). The ACE dev key rides on Authorization: Bearer — the provider's apiKey option sends it there and signs nothing — or on x-api-key, which the surface promotes to the bearer when none is present. That second shape is for a client that SigV4-signs against the gateway's host (accessKeyId / secretAccessKey set instead of apiKey): ACE never reads a credential out of a SigV4 Authorization and never forwards it — that signature was computed over the gateway's host and is useless upstream — so the ACE identity has to travel on x-api-key and the AWS credential on the x-ace-* headers. The streaming reply is AWS's binary event-stream framing, which is what a Converse client parses; Accept: application/x-ndjson asks for newline-delimited JSON instead.

import { createAmazonBedrock } from '@ai-sdk/amazon-bedrock';

const bedrock = createAmazonBedrock({
  baseURL: 'https://engine.acefleet.dev/bedrock',                 // the SDK builds /model/{modelId}/converse[-stream]
  apiKey: 'ace_dev_<your-dev-key>',                 // Authorization: Bearer — your ACE identity
  headers: {
    'x-ace-bedrock-api-key': 'ABSK...',       // zero-trust: sent upstream as a bearer, nothing signed
    'x-ace-aws-region': 'us-east-1',          // required with an API key
  },
});
// model: bedrock('anthropic.claude-3-5-sonnet-20241022-v2:0')
//
// With the SigV4 pair instead: drop apiKey and x-ace-bedrock-api-key, set
//   'x-api-key': 'ace_dev_<your-dev-key>',
//   'x-ace-bedrock-key': 'AKIAIOSFODNN7EXAMPLE',
//   'x-ace-aws-secret-access-key': 'wJalrXUtnFEMI/K7MDENG/...',
// in headers; the client-side signature the SDK adds is ignored and ACE signs for AWS itself.

Production hygiene. Put x-ace-mode: live in the factory's headers for production traffic: the per-request mode wins over the key's stored echo flag, so a key a test run left in echo cannot serve a synthetic 200 to real users. Send Cache-Control: no-cache on a call whose answer must not come from the semantic cache — a grounded search, a tool-bearing turn where freshness matters — through the call's own headers; the provider still answers and the answer is still stored, it is only never served from the cache (a request with temperature > 0.8 is not served from it either). Benchmark and load-test traffic against a production tenant carries x-ace-cache-partition, so it neither reads nor writes the pool that serves real users. To A/B one lever at a time from an unmodified client, send x-ace-skills (semantic_cache=off,prompt_compaction=shadow) on the call or the factory — the same override as /v1/execute's skill_overrides, nothing persisted — and read x-ace-skills-applied on the response for the pairs that actually ran; a skill your org has locked is a 403 in the provider's envelope, never a silent no-op. All four are in the request-header table.

import { generateText } from 'ai';
import { createAnthropic } from '@ai-sdk/anthropic';

const anthropic = createAnthropic({
  baseURL: 'https://engine.acefleet.dev/anthropic/v1',
  apiKey: 'ace_dev_<your-dev-key>',
  headers: { 'x-ace-anthropic-key': 'sk-ant-...', 'x-ace-mode': 'live' },
});

const { text } = await generateText({
  model: anthropic('claude-sonnet-4-5'),
  prompt: 'What changed in the incident channel in the last hour?',
  headers: {
    'cache-control': 'no-cache',                // this answer must be fresh
    'x-ace-skills': 'prompt_compaction=shadow', // measure the lever on this cohort, apply nothing
  },
});

Native response envelope

{
  "id": "req-a1b2c3",
  "version": "2026-08-01",
  "output": {
    "message": { "role": "assistant", "content": "Node 04 derated. Goodput restored." },
    "finish_reason": "stop"
  },
  "usage": {
    "input_tokens": 3980,
    "output_tokens": 212,
    "cost_usd": 0.0041,
    "cost_basis": "actual"
  },
  "trace": {
    "served_by": "openai:gpt-4o",
    "stages": [
      { "skill": "pii_ner", "mode": "prod", "action": "redacted", "kinds": { "EMAIL": 2 } },
      { "skill": "prompt_compaction", "mode": "shadow", "action": "would_compact", "would_save_usd": 0.0052 }
    ]
  }
}

The trace is the reason to prefer /v1/execute: it reports what every skill did, including skills running in shadow mode, which report a counterfactual (would_compact, would_save_usd) rather than changing the request.