Vercel AI SDK integration
Point each Vercel AI SDK provider (Anthropic, OpenAI, Azure, OpenRouter, Vertex, Bedrock) at ACE with baseURL and headers — no custom fetch.
An AI SDK provider is configured through the three knobs every createX() factory exposes — baseURL, headers and fetch — and on ACE the integration is the first two. baseURL moves to the gateway's vendor-shaped surface for that provider; the ACE developer key goes in apiKey, because each surface reads it from the header the provider already sends (x-api-key, api-key, x-goog-api-key or Authorization: Bearer); the upstream credential rides on the zero-trust header in headers. fetch is not needed on any provider: every path the six factories build, Bedrock's /model/{modelId}/converse[-stream] included, is one the gateway serves. The body the SDK serializes is what the vendor receives and the vendor's reply is what the SDK parses, streamed or not — the guarantees in Provider Surfaces apply unchanged.
Anthropic (@ai-sdk/anthropic)
apiKey is sent as x-api-key. The shim treats an ace_dev_ value there as your ACE identity and fills Authorization: Bearer with it when you set none; any other value is a pass-through Anthropic key. So the dev key goes in apiKey and the vendor key on x-ace-anthropic-key — or, equivalently, the vendor key in apiKey and Authorization: Bearer ace_dev_… in headers. Thinking, tools, images and cache_control are relayed as sent; the reply and its stream are Anthropic's own. Token counting is a plain POST of the same body and headers to /anthropic/v1/messages/count_tokens, which returns the vendor's count.
import { createAnthropic } from '@ai-sdk/anthropic';
const anthropic = createAnthropic({
baseURL: 'https://engine.acefleet.dev/anthropic/v1',
apiKey: 'ace_dev_<your-dev-key>', // sent as x-api-key: your ACE identity
headers: { 'x-ace-anthropic-key': 'sk-ant-...' }, // zero-trust: used for this call only
});
// model: anthropic('claude-sonnet-4-5')
OpenAI and Azure (@ai-sdk/openai, @ai-sdk/azure)
Both default to the Responses API, which the gateway serves as its own surface — /v1/responses, and under /azure/openai the three spellings the Azure SDKs use — with the item array, reasoning, store, include and encrypted_content relayed as sent and the response.* stream returned verbatim; chat completions from the same factories land on /v1/chat/completions and /azure/openai/deployments/{deployment}/chat/completions. @ai-sdk/openai sends apiKey as Authorization: Bearer; @ai-sdk/azure sends it as api-key, which the Azure surface reads the same way. Azure has no global host, so the resource endpoint travels on x-ace-azure-endpoint and the deployment name is the model.
import { createOpenAI } from '@ai-sdk/openai';
import { createAzure } from '@ai-sdk/azure';
const openai = createOpenAI({
baseURL: 'https://engine.acefleet.dev/v1',
apiKey: 'ace_dev_<your-dev-key>', // Authorization: Bearer
headers: { 'x-ace-openai-key': 'sk-proj-...' },
});
const azure = createAzure({
baseURL: 'https://engine.acefleet.dev/azure/openai',
apiKey: 'ace_dev_<your-dev-key>', // sent as api-key
headers: {
'x-ace-provider': 'azure',
'x-ace-azure-key': '<your-azure-key>',
'x-ace-azure-endpoint': 'https://<resource>.openai.azure.com',
},
});
// model: azure('gpt-4o-prod') — the deployment name
OpenRouter and other OpenAI-compatible hosts
Any OpenAI-compatible host is the openai channel pointed elsewhere: x-ace-openai-base-url names the host, x-ace-openai-key carries its key, and the model slug is sent verbatim — anthropic/claude-sonnet-4.5 reaches OpenRouter spelled exactly like that. x-ace-served-by still names the channel (openai), not the host. The same headers work from @openrouter/ai-sdk-provider or any provider package whose factory takes baseURL and headers, pointed at /v1. A Responses body on this channel is posted to {host}/responses, so a host that speaks only chat completions must be sent chat completions.
import { createOpenAI } from '@ai-sdk/openai';
const openrouter = createOpenAI({
baseURL: 'https://engine.acefleet.dev/v1',
apiKey: 'ace_dev_<your-dev-key>',
headers: {
'x-ace-provider': 'openai',
'x-ace-openai-base-url': 'https://openrouter.ai/api/v1',
'x-ace-openai-key': 'sk-or-v1-...',
},
});
// model: openrouter.chat('anthropic/claude-sonnet-4.5') — the host's own slug
Claude on Vertex, through the Anthropic shim
The one callers miss: no Vertex-specific package is needed. Claude on Vertex takes an Anthropic Messages body, which the Anthropic provider already produces, so the channel is only a different address and credential. Pinned with x-ace-provider: google_vertex, the same call goes to …/publishers/anthropic/models/{model}:rawPredict (:streamRawPredict when streaming): model moves out of the body into the path — use Vertex's id, e.g. claude-sonnet-4-5@20250929 — anthropic_version: vertex-2023-10-16 is set in the body, an anthropic-beta header travels as the anthropic_beta array, and the reply, thinking and signatures included, is Anthropic's verbatim. One of x-ace-vertex-token (a ya29 bearer) or x-ace-vertex-key (service-account JSON) is required, as is x-ace-google-project; x-ace-google-region defaults to us-central1.
import { createAnthropic } from '@ai-sdk/anthropic';
const claudeOnVertex = createAnthropic({
baseURL: 'https://engine.acefleet.dev/anthropic/v1',
apiKey: 'ace_dev_<your-dev-key>',
headers: {
'x-ace-provider': 'google_vertex',
'x-ace-vertex-token': 'ya29...', // or 'x-ace-vertex-key': '<service-account JSON>'
'x-ace-google-project': '<your-gcp-project>',
'x-ace-google-region': 'europe-west1',
},
});
// model: claudeOnVertex('claude-sonnet-4-5@20250929') — Vertex's id, sent in the path upstream
Gemini on Vertex, through the Gemini shim
/gemini/v1beta/models/{model}:generateContent (and :streamGenerateContent) takes Google's own GenerateContentRequest, served by the Gemini Developer API or — with x-ace-provider: google_vertex and the Vertex credential set above — by Vertex's publishers/google surface; the body goes upstream as written and the reply, thought parts and thoughtSignature included, comes back as sent. The credential contract is exact: the ACE dev key is read from Authorization: Bearer, or from x-goog-api-key when no bearer is present — a bearer you set yourself is never overwritten. A client that puts a Google OAuth token on Authorization: Bearer, the shape a Vertex-native client produces, therefore leaves the request with no ACE identity and is refused; whether a createVertex-built request can be made to fit is not something this reference claims. The Vertex token belongs on x-ace-vertex-token, and the dev key on x-goog-api-key (as @ai-sdk/google's createGoogleGenerativeAI sends apiKey) or on the bearer.
import { createGoogleGenerativeAI } from '@ai-sdk/google';
const geminiOnVertex = createGoogleGenerativeAI({
baseURL: 'https://engine.acefleet.dev/gemini/v1beta',
apiKey: 'ace_dev_<your-dev-key>', // sent as x-goog-api-key: your ACE identity
headers: {
'x-ace-provider': 'google_vertex',
'x-ace-vertex-token': 'ya29...', // or 'x-ace-vertex-key': '<service-account JSON>'
'x-ace-google-project': '<your-gcp-project>',
'x-ace-google-region': 'us-central1',
},
});
// Without x-ace-provider and the Vertex headers, the same client is served by the
// Gemini Developer API with 'x-ace-google-key': '<your-gemini-key>'.
Bedrock (@ai-sdk/amazon-bedrock)
baseURL is the gateway's /bedrock prefix and nothing else changes: the provider builds {baseURL}/model/{modelId}/converse[-stream] with the id percent-encoded in the path (: as %3A; an inference-profile ARN with / as %2F), and the gateway serves exactly that path, lifting modelId from it into the body. No request-rewriting fetch. The upstream credential is one of two on the zero-trust headers, never both: an Amazon Bedrock API key on x-ace-bedrock-api-key, which ACE sends upstream as a bearer with nothing signed, or the access-key pair on x-ace-bedrock-key + x-ace-aws-secret-access-key, which ACE SigV4-signs for the real AWS host. x-ace-aws-region names the region either way (required with the key). The ACE dev key rides on Authorization: Bearer — the provider's apiKey option sends it there and signs nothing — or on x-api-key, which the surface promotes to the bearer when none is present. That second shape is for a client that SigV4-signs against the gateway's host (accessKeyId / secretAccessKey set instead of apiKey): ACE never reads a credential out of a SigV4 Authorization and never forwards it — that signature was computed over the gateway's host and is useless upstream — so the ACE identity has to travel on x-api-key and the AWS credential on the x-ace-* headers. The streaming reply is AWS's binary event-stream framing, which is what a Converse client parses; Accept: application/x-ndjson asks for newline-delimited JSON instead.
import { createAmazonBedrock } from '@ai-sdk/amazon-bedrock';
const bedrock = createAmazonBedrock({
baseURL: 'https://engine.acefleet.dev/bedrock', // the SDK builds /model/{modelId}/converse[-stream]
apiKey: 'ace_dev_<your-dev-key>', // Authorization: Bearer — your ACE identity
headers: {
'x-ace-bedrock-api-key': 'ABSK...', // zero-trust: sent upstream as a bearer, nothing signed
'x-ace-aws-region': 'us-east-1', // required with an API key
},
});
// model: bedrock('anthropic.claude-3-5-sonnet-20241022-v2:0')
//
// With the SigV4 pair instead: drop apiKey and x-ace-bedrock-api-key, set
// 'x-api-key': 'ace_dev_<your-dev-key>',
// 'x-ace-bedrock-key': 'AKIAIOSFODNN7EXAMPLE',
// 'x-ace-aws-secret-access-key': 'wJalrXUtnFEMI/K7MDENG/...',
// in headers; the client-side signature the SDK adds is ignored and ACE signs for AWS itself.
Production hygiene
Put x-ace-mode: live in the factory's headers for production traffic: the per-request mode wins over the key's stored echo flag, so a key a test run left in echo cannot serve a synthetic 200 to real users. Send Cache-Control: no-cache on a call whose answer must not come from the semantic cache — a grounded search, a tool-bearing turn where freshness matters — through the call's own headers; the provider still answers and the answer is still stored, it is only never served from the cache (a request with temperature > 0.8 is not served from it either). Benchmark and load-test traffic against a production tenant carries x-ace-cache-partition, so it neither reads nor writes the pool that serves real users. To A/B one lever at a time from an unmodified client, send x-ace-skills (semantic_cache=off,prompt_compaction=shadow) on the call or the factory — the same override as /v1/execute's skill_overrides, nothing persisted — and read x-ace-skills-applied on the response for the pairs that actually ran; a skill your org has locked is a 403 in the provider's envelope, never a silent no-op. All four are in the request-header table.
import { generateText } from 'ai';
import { createAnthropic } from '@ai-sdk/anthropic';
const anthropic = createAnthropic({
baseURL: 'https://engine.acefleet.dev/anthropic/v1',
apiKey: 'ace_dev_<your-dev-key>',
headers: { 'x-ace-anthropic-key': 'sk-ant-...', 'x-ace-mode': 'live' },
});
const { text } = await generateText({
model: anthropic('claude-sonnet-4-5'),
prompt: 'What changed in the incident channel in the last hour?',
headers: {
'cache-control': 'no-cache', // this answer must be fresh
'x-ace-skills': 'prompt_compaction=shadow', // measure the lever on this cohort, apply nothing
},
});