all docs
/ integrations

LiteLLM integration

Run LiteLLM traffic through the ACE gateway: endpoint, the headers a call sets, what passes through untouched, and a runnable example.

Put ACE behind LiteLLM (client → LiteLLM → ACE) for one admin change and no app change, or in front of it (client → ACE → LiteLLM) for full coverage. LiteLLM's routing, budgets, virtual keys, retries and fallbacks stay in charge; ACE makes one attempt.

  • ACE behind: each model_list entry gets api_base = ACE (/v1 for openai/ models, /anthropic for anthropic/) and api_key = the ACE key; provider keys are stored on ACE.
  • Spend: ACE returns the provider's usage verbatim, cache tokens included, so x-litellm-response-cost and spend logs price each call as the provider would.
  • Caveat: a LiteLLM fallback to a target that is not ACE bypasses ACE.

Endpoint: https://engine.acefleet.dev/v1/chat/completions · also https://engine.acefleet.dev/v1/responses, https://engine.acefleet.dev/anthropic/v1/messages. ACE credential: Authorization: Bearer ace_dev_... (x-api-key on /anthropic/v1/messages).

x-ace-upstream-profile: litellm names the profile.

Topology Shape
ACE in front
Full coverage. Client → ACE → your LiteLLM proxy. The proxy must be a public https host for the managed gateway; on-prem may allowlist an internal one.
ACE behind
Default. Client → LiteLLM → ACE: api_base is ACE's /v1 (OpenAI models, sent as Authorization: Bearer) or /anthropic (Anthropic models, sent as x-api-key), api_key is the ACE key, provider keys are stored on ACE, and forward_client_headers_to_llm_api: true carries x-ace-* headers through — every client x-* header, so callers can steer ACE's mode, provider and upstream; a stored provider key is still never sent to a host they chose (403). LiteLLM's /anthropic pass-through takes ANTHROPIC_API_BASE and ANTHROPIC_API_KEY instead of model_list. An admin-only change; a fallback to a non-ACE target bypasses ACE.
Header Required Description
x-ace-openai-base-url
Yes
The LiteLLM proxy URL, e.g. https://litellm.example.com/v1. Not needed when stored as the endpoint of the tenant's openai vendor key.
x-ace-openai-key
Yes
A LiteLLM virtual key. Not needed when stored as the tenant's openai vendor key. On /anthropic/v1/messages it rides on x-ace-anthropic-key instead.

Passed through untouched

  • x-litellm-* request headers (-tags, -spend-logs-metadata, -customer-id, -end-user-id, -trace-id, -session-id, -num-retries, -timeout)
  • body fields metadata, litellm_metadata, tags, user, cache, fallbacks, num_retries, guardrails
  • model group names, verbatim
  • x-litellm-* response headers and : ping stream frames

Storing the virtual key, the proxy URL as its endpoint and the litellm profile as the tenant's openai (or anthropic) vendor key leaves the client changing only base_url and api_key; no x-ace-provider is needed. The tenant has one vendor key per channel, so this replaces any provider key stored there and sends all of the tenant's traffic on that channel through LiteLLM.

LiteLLM with ACE (config.yaml + cURL)

# ACE behind LiteLLM: each model's api_base is ACE; provider keys are stored on ACE.
cat > config.yaml <<'EOF'
model_list:
  - model_name: gpt-via-ace
    litellm_params:
      model: openai/gpt-4o
      api_base: https://engine.acefleet.dev/v1               # LiteLLM sends Authorization: Bearer $ACE_KEY
      api_key: os.environ/ACE_KEY
  - model_name: claude-via-ace
    litellm_params:
      model: anthropic/claude-sonnet-4-5
      api_base: https://engine.acefleet.dev/anthropic        # LiteLLM sends x-api-key: $ACE_KEY to /anthropic/v1/messages
      api_key: os.environ/ACE_KEY
general_settings:
  forward_client_headers_to_llm_api: true   # forwards every client x-* header, e.g. x-ace-session
EOF
export ACE_KEY="ace_dev_<your-dev-key>"
# LiteLLM's /anthropic pass-through route ignores model_list; it uses these two.
export ANTHROPIC_API_BASE="https://engine.acefleet.dev/anthropic"
export ANTHROPIC_API_KEY="$ACE_KEY"
litellm --config config.yaml

curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Authorization: Bearer sk-<litellm-virtual-key>" \
  -H "x-ace-session: nightly-triage-0930" \
  -H "Content-Type: application/json" \
  -d '{ "model": "gpt-via-ace", "messages": [{ "role": "user", "content": "Summarise the failing checks." }] }'

# The same proxy on LiteLLM's Anthropic-shaped route.
curl -X POST http://localhost:4000/v1/messages \
  -H "x-api-key: sk-<litellm-virtual-key>" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{ "model": "claude-via-ace", "max_tokens": 1024, "messages": [{ "role": "user", "content": "Summarise the failing checks." }] }'

# ACE in front of LiteLLM, per request.
curl -X POST https://engine.acefleet.dev/v1/chat/completions \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "x-ace-upstream-profile: litellm" \
  -H "x-ace-openai-base-url: https://litellm.example.com/v1" \
  -H "x-ace-openai-key: sk-<litellm-virtual-key>" \
  -H "x-litellm-tags: nightly,triage" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{ "role": "user", "content": "Summarise the failing checks." }],
    "metadata": { "trace_name": "nightly-triage" }
  }'

# ACE in front of LiteLLM, stored once as the tenant's openai vendor key: the proxy URL, a
# virtual key and the litellm profile.
curl -X POST https://engine.acefleet.dev/api/v1/vendor_key/create \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "Content-Type: application/json" \
  -d '{ "provider": "openai", "api_key": "sk-<litellm-virtual-key>", "endpoint": "https://litellm.example.com/v1", "upstream_profile": "litellm" }'

# Then only the base URL and the API key change on the client.
curl -X POST https://engine.acefleet.dev/v1/chat/completions \
  -H "Authorization: Bearer ace_dev_<your-dev-key>" \
  -H "Content-Type: application/json" \
  -d '{ "model": "gpt-4o", "messages": [{ "role": "user", "content": "Summarise the failing checks." }] }'