Integration Quickstarts
Copy-pasteable OpenAI (chat completions and the Responses API), Anthropic, Azure, Gemini, Vertex, Bedrock and native-engine integrations, and the Vercel AI SDK providers (including Claude on Vertex through the Anthropic provider) — the drop-in is a base URL change plus one header.
ACE is a drop-in reverse proxy. Point a standard SDK at the gateway with your ACE developer key and pass the upstream provider credential as a header.
OpenAI SDK (Python)
import openai
client = openai.OpenAI(
base_url="https://engine.acefleet.dev/v1",
api_key="ace_dev_<your-dev-key>",
default_headers={
"x-ace-openai-key": "sk-proj-...", # Zero-Trust header mode
},
)
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Analyze cluster telemetry."}],
)
print(response.choices[0].message.content)
OpenAI proxy (cURL)
curl -X POST https://engine.acefleet.dev/v1/chat/completions \
-H "Authorization: Bearer ace_dev_<your-dev-key>" \
-H "x-ace-openai-key: sk-proj-..." \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{ "role": "user", "content": "Analyze cluster telemetry." }]
}'
OpenAI Responses API (Python)
from openai import OpenAI
client = OpenAI(
base_url="https://engine.acefleet.dev/v1",
api_key="ace_dev_<your-dev-key>",
default_headers={
"x-ace-openai-key": "sk-proj-...", # Zero-Trust header mode
},
)
# Item-based, not message-based. store=False + include=[...] keeps reasoning
# continuity across turns without OpenAI retaining the conversation.
response = client.responses.create(
model="gpt-5",
instructions="You are a concise assistant.",
input=[{"role": "user", "content": [{"type": "input_text", "text": "Analyze cluster telemetry."}]}],
reasoning={"effort": "medium", "summary": "auto"},
store=False,
include=["reasoning.encrypted_content"],
)
print(response.output_text)
Azure Responses API (cURL)
curl -X POST "https://engine.acefleet.dev/azure/openai/v1/responses?api-version=preview" \
-H "Authorization: Bearer ace_dev_<your-dev-key>" \
-H "x-ace-provider: azure" \
-H "x-ace-azure-key: <your-azure-key>" \
-H "x-ace-azure-endpoint: https://<resource>.openai.azure.com" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5-mini",
"input": [{ "role": "user", "content": [{ "type": "input_text", "text": "Optimize GPU allocations." }] }],
"reasoning": { "effort": "low" },
"store": false,
"include": ["reasoning.encrypted_content"]
}'
Anthropic proxy (cURL)
curl -X POST https://engine.acefleet.dev/anthropic/v1/messages \
-H "Authorization: Bearer ace_dev_<your-dev-key>" \
-H "x-ace-anthropic-key: sk-ant-..." \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "Optimize GPU allocations." }]
}'
Azure OpenAI proxy (cURL)
curl -X POST "https://engine.acefleet.dev/azure/openai/deployments/gpt-4o-prod/chat/completions?api-version=2024-02-15-preview" \
-H "Authorization: Bearer ace_dev_<your-dev-key>" \
-H "x-ace-provider: azure" \
-H "x-ace-azure-key: <your-azure-key>" \
-H "x-ace-azure-endpoint: https://<resource>.openai.azure.com" \
-H "Content-Type: application/json" \
-d '{ "messages": [{ "role": "user", "content": "Optimize GPU allocations." }] }'
Google Gemini proxy (cURL)
curl -X POST "https://engine.acefleet.dev/gemini/v1beta/models/gemini-2.5-flash:generateContent" \
-H "x-goog-api-key: ace_dev_<your-dev-key>" \
-H "x-ace-provider: google_gemini" \
-H "x-ace-google-key: <your-gemini-key>" \
-H "Content-Type: application/json" \
-d '{ "contents": [{ "role": "user", "parts": [{ "text": "Optimize GPU allocations." }] }] }'
Google Vertex AI proxy (cURL)
curl -X POST "https://engine.acefleet.dev/gemini/v1beta/models/gemini-2.5-flash:generateContent" \
-H "x-goog-api-key: ace_dev_<your-dev-key>" \
-H "x-ace-provider: google_vertex" \
-H "x-ace-vertex-token: ya29..." \
-H "x-ace-google-project: <your-gcp-project>" \
-H "x-ace-google-region: us-central1" \
-H "Content-Type: application/json" \
-d '{ "contents": [{ "role": "user", "parts": [{ "text": "Optimize GPU allocations." }] }] }'
AWS Bedrock proxy (cURL)
# An Amazon Bedrock API key: sent upstream as a bearer, nothing signed, no AWS secret
# leaves your side. The region is required with it.
curl -X POST https://engine.acefleet.dev/bedrock/converse \
-H "Authorization: Bearer ace_dev_<your-dev-key>" \
-H "x-ace-bedrock-api-key: ABSK..." \
-H "x-ace-aws-region: us-east-1" \
-H "Content-Type: application/json" \
-d '{
"modelId": "anthropic.claude-3-5-sonnet-20241022-v2:0",
"messages": [
{ "role": "user", "content": [{ "text": "Optimize GPU allocations." }] }
]
}'
# The SigV4 pair instead (never both): ACE signs the request for the AWS host.
# modelId may also travel in the path, as an AWS SDK sends it (percent-encoded).
curl -X POST https://engine.acefleet.dev/bedrock/model/anthropic.claude-3-5-sonnet-20241022-v2%3A0/converse \
-H "Authorization: Bearer ace_dev_<your-dev-key>" \
-H "x-ace-bedrock-key: AKIAIOSFODNN7EXAMPLE" \
-H "x-ace-aws-secret-access-key: wJalrXUtnFEMI/K7MDENG/..." \
-H "x-ace-aws-region: us-east-1" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "user", "content": [{ "text": "Optimize GPU allocations." }] }
]
}'
ACE native engine (cURL)
curl -X POST https://engine.acefleet.dev/v1/execute \
-H "Authorization: Bearer ace_dev_<your-dev-key>" \
-H "Content-Type: application/json" \
-d '{
"input": { "messages": [{ "role": "user", "content": "Audit node health." }] },
"target": { "model": "gpt-4o" },
"execution": { "trace": "full" }
}'
Vercel AI SDK providers
An AI SDK provider is configured through the three knobs every createX() factory exposes — baseURL, headers and fetch — and on ACE the integration is the first two. baseURL moves to the gateway's vendor-shaped surface for that provider; the ACE developer key goes in apiKey, because each surface reads it from the header the provider already sends (x-api-key, api-key, x-goog-api-key or Authorization: Bearer); the upstream credential rides on the zero-trust header in headers. fetch is not needed on any provider: every path the six factories build, Bedrock's /model/{modelId}/converse[-stream] included, is one the gateway serves. The body the SDK serializes is what the vendor receives and the vendor's reply is what the SDK parses, streamed or not — the guarantees in Provider Surfaces apply unchanged.
Anthropic (@ai-sdk/anthropic). apiKey is sent as x-api-key. The shim treats an ace_dev_ value there as your ACE identity and fills Authorization: Bearer with it when you set none; any other value is a pass-through Anthropic key. So the dev key goes in apiKey and the vendor key on x-ace-anthropic-key — or, equivalently, the vendor key in apiKey and Authorization: Bearer ace_dev_… in headers. Thinking, tools, images and cache_control are relayed as sent; the reply and its stream are Anthropic's own. Token counting is a plain POST of the same body and headers to /anthropic/v1/messages/count_tokens, which returns the vendor's count.
import { createAnthropic } from '@ai-sdk/anthropic';
const anthropic = createAnthropic({
baseURL: 'https://engine.acefleet.dev/anthropic/v1',
apiKey: 'ace_dev_<your-dev-key>', // sent as x-api-key: your ACE identity
headers: { 'x-ace-anthropic-key': 'sk-ant-...' }, // zero-trust: used for this call only
});
// model: anthropic('claude-sonnet-4-5')
OpenAI and Azure (@ai-sdk/openai, @ai-sdk/azure). Both default to the Responses API, which the gateway serves as its own surface — /v1/responses, and under /azure/openai the three spellings the Azure SDKs use — with the item array, reasoning, store, include and encrypted_content relayed as sent and the response.* stream returned verbatim; chat completions from the same factories land on /v1/chat/completions and /azure/openai/deployments/{deployment}/chat/completions. @ai-sdk/openai sends apiKey as Authorization: Bearer; @ai-sdk/azure sends it as api-key, which the Azure surface reads the same way. Azure has no global host, so the resource endpoint travels on x-ace-azure-endpoint and the deployment name is the model.
import { createOpenAI } from '@ai-sdk/openai';
import { createAzure } from '@ai-sdk/azure';
const openai = createOpenAI({
baseURL: 'https://engine.acefleet.dev/v1',
apiKey: 'ace_dev_<your-dev-key>', // Authorization: Bearer
headers: { 'x-ace-openai-key': 'sk-proj-...' },
});
const azure = createAzure({
baseURL: 'https://engine.acefleet.dev/azure/openai',
apiKey: 'ace_dev_<your-dev-key>', // sent as api-key
headers: {
'x-ace-provider': 'azure',
'x-ace-azure-key': '<your-azure-key>',
'x-ace-azure-endpoint': 'https://<resource>.openai.azure.com',
},
});
// model: azure('gpt-4o-prod') — the deployment name
OpenRouter and other OpenAI-compatible hosts. Any OpenAI-compatible host is the openai channel pointed elsewhere: x-ace-openai-base-url names the host, x-ace-openai-key carries its key, and the model slug is sent verbatim — anthropic/claude-sonnet-4.5 reaches OpenRouter spelled exactly like that. x-ace-served-by still names the channel (openai), not the host. The same headers work from @openrouter/ai-sdk-provider or any provider package whose factory takes baseURL and headers, pointed at /v1. A Responses body on this channel is posted to {host}/responses, so a host that speaks only chat completions must be sent chat completions.
import { createOpenAI } from '@ai-sdk/openai';
const openrouter = createOpenAI({
baseURL: 'https://engine.acefleet.dev/v1',
apiKey: 'ace_dev_<your-dev-key>',
headers: {
'x-ace-provider': 'openai',
'x-ace-openai-base-url': 'https://openrouter.ai/api/v1',
'x-ace-openai-key': 'sk-or-v1-...',
},
});
// model: openrouter.chat('anthropic/claude-sonnet-4.5') — the host's own slug
Claude on Vertex, through the Anthropic shim. The one callers miss: no Vertex-specific package is needed. Claude on Vertex takes an Anthropic Messages body, which the Anthropic provider already produces, so the channel is only a different address and credential. Pinned with x-ace-provider: google_vertex, the same call goes to …/publishers/anthropic/models/{model}:rawPredict (:streamRawPredict when streaming): model moves out of the body into the path — use Vertex's id, e.g. claude-sonnet-4-5@20250929 — anthropic_version: vertex-2023-10-16 is set in the body, an anthropic-beta header travels as the anthropic_beta array, and the reply, thinking and signatures included, is Anthropic's verbatim. One of x-ace-vertex-token (a ya29 bearer) or x-ace-vertex-key (service-account JSON) is required, as is x-ace-google-project; x-ace-google-region defaults to us-central1.
import { createAnthropic } from '@ai-sdk/anthropic';
const claudeOnVertex = createAnthropic({
baseURL: 'https://engine.acefleet.dev/anthropic/v1',
apiKey: 'ace_dev_<your-dev-key>',
headers: {
'x-ace-provider': 'google_vertex',
'x-ace-vertex-token': 'ya29...', // or 'x-ace-vertex-key': '<service-account JSON>'
'x-ace-google-project': '<your-gcp-project>',
'x-ace-google-region': 'europe-west1',
},
});
// model: claudeOnVertex('claude-sonnet-4-5@20250929') — Vertex's id, sent in the path upstream
Gemini on Vertex, through the Gemini shim. /gemini/v1beta/models/{model}:generateContent (and :streamGenerateContent) takes Google's own GenerateContentRequest, served by the Gemini Developer API or — with x-ace-provider: google_vertex and the Vertex credential set above — by Vertex's publishers/google surface; the body goes upstream as written and the reply, thought parts and thoughtSignature included, comes back as sent. The credential contract is exact: the ACE dev key is read from Authorization: Bearer, or from x-goog-api-key when no bearer is present — a bearer you set yourself is never overwritten. A client that puts a Google OAuth token on Authorization: Bearer, the shape a Vertex-native client produces, therefore leaves the request with no ACE identity and is refused; whether a createVertex-built request can be made to fit is not something this reference claims. The Vertex token belongs on x-ace-vertex-token, and the dev key on x-goog-api-key (as @ai-sdk/google's createGoogleGenerativeAI sends apiKey) or on the bearer.
import { createGoogleGenerativeAI } from '@ai-sdk/google';
const geminiOnVertex = createGoogleGenerativeAI({
baseURL: 'https://engine.acefleet.dev/gemini/v1beta',
apiKey: 'ace_dev_<your-dev-key>', // sent as x-goog-api-key: your ACE identity
headers: {
'x-ace-provider': 'google_vertex',
'x-ace-vertex-token': 'ya29...', // or 'x-ace-vertex-key': '<service-account JSON>'
'x-ace-google-project': '<your-gcp-project>',
'x-ace-google-region': 'us-central1',
},
});
// Without x-ace-provider and the Vertex headers, the same client is served by the
// Gemini Developer API with 'x-ace-google-key': '<your-gemini-key>'.
Bedrock (@ai-sdk/amazon-bedrock). baseURL is the gateway's /bedrock prefix and nothing else changes: the provider builds {baseURL}/model/{modelId}/converse[-stream] with the id percent-encoded in the path (: as %3A; an inference-profile ARN with / as %2F), and the gateway serves exactly that path, lifting modelId from it into the body. No request-rewriting fetch. The upstream credential is one of two on the zero-trust headers, never both: an Amazon Bedrock API key on x-ace-bedrock-api-key, which ACE sends upstream as a bearer with nothing signed, or the access-key pair on x-ace-bedrock-key + x-ace-aws-secret-access-key, which ACE SigV4-signs for the real AWS host. x-ace-aws-region names the region either way (required with the key). The ACE dev key rides on Authorization: Bearer — the provider's apiKey option sends it there and signs nothing — or on x-api-key, which the surface promotes to the bearer when none is present. That second shape is for a client that SigV4-signs against the gateway's host (accessKeyId / secretAccessKey set instead of apiKey): ACE never reads a credential out of a SigV4 Authorization and never forwards it — that signature was computed over the gateway's host and is useless upstream — so the ACE identity has to travel on x-api-key and the AWS credential on the x-ace-* headers. The streaming reply is AWS's binary event-stream framing, which is what a Converse client parses; Accept: application/x-ndjson asks for newline-delimited JSON instead.
import { createAmazonBedrock } from '@ai-sdk/amazon-bedrock';
const bedrock = createAmazonBedrock({
baseURL: 'https://engine.acefleet.dev/bedrock', // the SDK builds /model/{modelId}/converse[-stream]
apiKey: 'ace_dev_<your-dev-key>', // Authorization: Bearer — your ACE identity
headers: {
'x-ace-bedrock-api-key': 'ABSK...', // zero-trust: sent upstream as a bearer, nothing signed
'x-ace-aws-region': 'us-east-1', // required with an API key
},
});
// model: bedrock('anthropic.claude-3-5-sonnet-20241022-v2:0')
//
// With the SigV4 pair instead: drop apiKey and x-ace-bedrock-api-key, set
// 'x-api-key': 'ace_dev_<your-dev-key>',
// 'x-ace-bedrock-key': 'AKIAIOSFODNN7EXAMPLE',
// 'x-ace-aws-secret-access-key': 'wJalrXUtnFEMI/K7MDENG/...',
// in headers; the client-side signature the SDK adds is ignored and ACE signs for AWS itself.
Production hygiene. Put x-ace-mode: live in the factory's headers for production traffic: the per-request mode wins over the key's stored echo flag, so a key a test run left in echo cannot serve a synthetic 200 to real users. Send Cache-Control: no-cache on a call whose answer must not come from the semantic cache — a grounded search, a tool-bearing turn where freshness matters — through the call's own headers; the provider still answers and the answer is still stored, it is only never served from the cache (a request with temperature > 0.8 is not served from it either). Benchmark and load-test traffic against a production tenant carries x-ace-cache-partition, so it neither reads nor writes the pool that serves real users. To A/B one lever at a time from an unmodified client, send x-ace-skills (semantic_cache=off,prompt_compaction=shadow) on the call or the factory — the same override as /v1/execute's skill_overrides, nothing persisted — and read x-ace-skills-applied on the response for the pairs that actually ran; a skill your org has locked is a 403 in the provider's envelope, never a silent no-op. All four are in the request-header table.
import { generateText } from 'ai';
import { createAnthropic } from '@ai-sdk/anthropic';
const anthropic = createAnthropic({
baseURL: 'https://engine.acefleet.dev/anthropic/v1',
apiKey: 'ace_dev_<your-dev-key>',
headers: { 'x-ace-anthropic-key': 'sk-ant-...', 'x-ace-mode': 'live' },
});
const { text } = await generateText({
model: anthropic('claude-sonnet-4-5'),
prompt: 'What changed in the incident channel in the last hour?',
headers: {
'cache-control': 'no-cache', // this answer must be fresh
'x-ace-skills': 'prompt_compaction=shadow', // measure the lever on this cohort, apply nothing
},
});
Native response envelope
{
"id": "req-a1b2c3",
"version": "2026-08-01",
"output": {
"message": { "role": "assistant", "content": "Node 04 derated. Goodput restored." },
"finish_reason": "stop"
},
"usage": {
"input_tokens": 3980,
"output_tokens": 212,
"cost_usd": 0.0041,
"cost_basis": "actual"
},
"trace": {
"served_by": "openai:gpt-4o",
"stages": [
{ "skill": "pii_ner", "mode": "prod", "action": "redacted", "kinds": { "EMAIL": 2 } },
{ "skill": "prompt_compaction", "mode": "shadow", "action": "would_compact", "would_save_usd": 0.0052 }
]
}
}
The trace is the reason to prefer /v1/execute: it reports what every skill did, including skills running in shadow mode, which report a counterfactual (would_compact, would_save_usd) rather than changing the request.