Google Vertex AI integration
Call Google Vertex AI through the ACE gateway: endpoint, zero-trust credential headers, stored-key alternative and runnable examples.
Send Google Vertex AI traffic to https://engine.acefleet.dev/gemini/v1beta/models/<model>:generateContent with your ACE developer key in Authorization: Bearer ace_dev_... and the provider credential on the zero-trust headers below. The request body is the vendor's own, relayed as written, and the vendor's answer comes back verbatim.
The same headers also serve:
https://engine.acefleet.dev/gemini/v1beta/models/<model>:streamGenerateContenthttps://engine.acefleet.dev/anthropic/v1/messageshttps://engine.acefleet.dev/v1/chat/completions
Zero-trust headers
| Header | Required | What it carries |
|---|---|---|
x-ace-provider |
yes | google_vertex. Unpinned, the Gemini-shaped path is served by the Gemini Developer API and a stored Vertex key is never consulted. |
x-ace-vertex-key |
situational | GCP service-account JSON (raw or Base64). Use this OR x-ace-vertex-token, not both — the value is parsed as JSON, so a ya29 token here fails as malformed. |
x-ace-vertex-token |
situational | Direct OAuth bearer (ya29…), e.g. from gcloud auth print-access-token. The alternative to x-ace-vertex-key; one of the two is required. |
x-ace-google-project |
yes | GCP project id. Auto-extracted when passing service-account JSON. |
x-ace-google-region |
situational | Regional publisher endpoint. Defaults to us-central1; global is addressed at the unprefixed aiplatform.googleapis.com host. |
Stored key instead
A stored google_vertex vendor key replaces the credential header; project and region are stored with it — x-ace-provider: google_vertex is still required. Gemini models on Vertex take Google's own GenerateContentRequest on /gemini/v1beta/models/{model}:generateContent (:streamGenerateContent when streaming): the body goes to publishers/google/models/{model} as written and the reply, thought parts and thoughtSignature included, comes back as sent. There is no /v1/projects/{project}/locations/{location}/… route on the gateway; project and region are headers, and a Google OAuth token belongs on x-ace-vertex-token, never on Authorization (that header carries the ACE key). An OpenAI-shaped body on /v1/chat/completions pinned to this channel is translated to the same upstream and answered in OpenAI's shape. Claude models on Vertex are served from /anthropic/v1/messages with the same credential and x-ace-provider: google_vertex: your Messages body goes to publishers/anthropic/models/{model}:rawPredict (:streamRawPredict when streaming) spelled as Anthropic's own Vertex SDK spells it — model in the path (use Vertex's id, e.g. claude-sonnet-4-5@20250929), anthropic_version: vertex-2023-10-16 in the body, your anthropic-beta header as the anthropic_beta array — and the reply, thinking and signatures included, is the vendor's verbatim. A Claude model pinned to this channel on /v1/chat/completions is refused: that surface has no Anthropic body to send. From an AI SDK client this is createAnthropic with these headers and no Vertex package — see the Vercel AI SDK quickstart.
Example request
curl -X POST "https://engine.acefleet.dev/gemini/v1beta/models/gemini-2.5-flash:generateContent" \
-H "x-goog-api-key: ace_dev_<your-dev-key>" \
-H "x-ace-provider: google_vertex" \
-H "x-ace-vertex-token: ya29..." \
-H "x-ace-google-project: <your-gcp-project>" \
-H "x-ace-google-region: us-central1" \
-H "Content-Type: application/json" \
-d '{ "contents": [{ "role": "user", "parts": [{ "text": "ping" }] }] }'
# Claude on Vertex, from the Anthropic surface (EU residency: region europe-west1):
curl -X POST "https://engine.acefleet.dev/anthropic/v1/messages" \
-H "x-api-key: ace_dev_<your-dev-key>" \
-H "x-ace-provider: google_vertex" \
-H "x-ace-vertex-key: <base64-service-account-json>" \
-H "x-ace-google-project: <your-gcp-project>" \
-H "x-ace-google-region: europe-west1" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{ "model": "claude-sonnet-4-5@20250929", "max_tokens": 1024, "messages": [{ "role": "user", "content": "ping" }] }'
Quickstarts
Google Vertex AI proxy (cURL)
curl -X POST "https://engine.acefleet.dev/gemini/v1beta/models/gemini-2.5-flash:generateContent" \
-H "x-goog-api-key: ace_dev_<your-dev-key>" \
-H "x-ace-provider: google_vertex" \
-H "x-ace-vertex-token: ya29..." \
-H "x-ace-google-project: <your-gcp-project>" \
-H "x-ace-google-region: us-central1" \
-H "Content-Type: application/json" \
-d '{ "contents": [{ "role": "user", "parts": [{ "text": "Optimize GPU allocations." }] }] }'
Vercel AI SDK: Claude on Vertex, through the Anthropic shim
The one callers miss: no Vertex-specific package is needed. Claude on Vertex takes an Anthropic Messages body, which the Anthropic provider already produces, so the channel is only a different address and credential. Pinned with x-ace-provider: google_vertex, the same call goes to …/publishers/anthropic/models/{model}:rawPredict (:streamRawPredict when streaming): model moves out of the body into the path — use Vertex's id, e.g. claude-sonnet-4-5@20250929 — anthropic_version: vertex-2023-10-16 is set in the body, an anthropic-beta header travels as the anthropic_beta array, and the reply, thinking and signatures included, is Anthropic's verbatim. One of x-ace-vertex-token (a ya29 bearer) or x-ace-vertex-key (service-account JSON) is required, as is x-ace-google-project; x-ace-google-region defaults to us-central1.
import { createAnthropic } from '@ai-sdk/anthropic';
const claudeOnVertex = createAnthropic({
baseURL: 'https://engine.acefleet.dev/anthropic/v1',
apiKey: 'ace_dev_<your-dev-key>',
headers: {
'x-ace-provider': 'google_vertex',
'x-ace-vertex-token': 'ya29...', // or 'x-ace-vertex-key': '<service-account JSON>'
'x-ace-google-project': '<your-gcp-project>',
'x-ace-google-region': 'europe-west1',
},
});
// model: claudeOnVertex('claude-sonnet-4-5@20250929') — Vertex's id, sent in the path upstream