AI inference fabric / online

The control plane for AI compute.

Route less. Reuse more. Give every request the right model, context, and compute.

no application rewrite OpenAI · Azure · Anthropic · neoclouds · your GPUs
signal / 0048adaptive runtime
ACE / AI-COMPUTE-EFFICIENCYscroll to stress the system
select a starting stack01 / 03
selected / managed apis

Managed Cloud APIs

60s Setup (OpenAI, Azure, Anthropic). One integration surface, no application rewrite.

start here
ace_efficiency_engine · projected live

→ Every prompt cached, routed, and pruned before it burns a token.

prompt_tokens−58%
cache_hit_ratio0.71
avg_latency18ms
cost_per_1k_calls$0.34
/ already on openai or anthropic?

Change one line. Your SDK calls, streaming and tool use stay exactly as they are.

python
base_url = "https://engine.acefleet.dev/v1"
  • SOC2 Type II
  • Zero data retention
  • Local ONNX embeddings
  • Self-hosted VPC
  • K8s operator native
✳route intent✳compress context✳reuse tokens✳allocate compute✳observe quietly✳route intent✳compress context✳reuse tokens✳allocate compute✳observe quietly
/ the problem

Your cloud API bill is growing
faster than your product.

Every layer of your AI stack — from silicon to agent frameworks — adds cost, latency, and fragility. ACE operates across all of them.

HWHardware & Accelerators·NVIDIA H100/H200 · Groq LPU · Cerebras WSE-3 · AMD Instinct
ENGInference Engines & Serving·vLLM · SGLang · TensorRT-LLM · TGI
CLDClouds & Orchestration·AWS Bedrock · Azure · GCP · CoreWeave · Kubernetes · Ray
AGTAgent Frameworks·CrewAI · LangChain · LlamaIndex · DSPy
HWSilicon & Accelerators·NVIDIA H100/H200 · Groq LPU · Cerebras WSE-3 · AMD Instinct
NETInterconnect·NVLink · InfiniBand · RoCEv2
CLDNeoclouds & Hyperscalers·CoreWeave · Nebius · Lambda · AWS · GCP · Azure
ORCOrchestration·Kubernetes · Ray · Karpenter
SRVInference Engines·vLLM · SGLang · TensorRT-LLM · TGI
OSSCustom Open-Source Models·Llama · Qwen · DeepSeek · Mistral
ENTEnterprise Platforms·Databricks · Snowflake · Vertex · Bedrock
AGTFrontier APIs & Agents·OpenAI · Anthropic · CrewAI · LangGraph · LlamaIndex · DSPy
HWHardware & Accelerators·NVIDIA H100/H200 · Groq LPU · Cerebras WSE-3 · AMD Instinct
ENGInference Engines & Serving·vLLM · SGLang · TensorRT-LLM · TGI
CLDClouds & Orchestration·AWS Bedrock · Azure · GCP · CoreWeave · Kubernetes · Ray
AGTAgent Frameworks·CrewAI · LangChain · LlamaIndex · DSPy
HWSilicon & Accelerators·NVIDIA H100/H200 · Groq LPU · Cerebras WSE-3 · AMD Instinct
NETInterconnect·NVLink · InfiniBand · RoCEv2
CLDNeoclouds & Hyperscalers·CoreWeave · Nebius · Lambda · AWS · GCP · Azure
ORCOrchestration·Kubernetes · Ray · Karpenter
SRVInference Engines·vLLM · SGLang · TensorRT-LLM · TGI
OSSCustom Open-Source Models·Llama · Qwen · DeepSeek · Mistral
ENTEnterprise Platforms·Databricks · Snowflake · Vertex · Bedrock
AGTFrontier APIs & Agents·OpenAI · Anthropic · CrewAI · LangGraph · LlamaIndex · DSPy
/ savings instrument

Find the work
you can give back.

Drop in your monthly spend and see what ACE would have saved on it. Cascaded waterfall math — no overlapping percentages, no black-box multipliers.

01 monthly AI spend $100,000
$1k$500k+
est. P95 latency reduction20%
02 workload profile
03 active optimization skills
projected reclaim

46%

of metered spend, modelled on 30% cacheable · 30% routable · 30% compressible traffic.

modelled annual savings · expected case$546,329conservative $382,430 · aggressive $655,595
current monthly $100,000optimized monthly $54,473monthly savings $45,527
04 where the spend goes
cache router prune net spend
verify this on your own traffic zero code refactoring · 60-second setup
/ live fabric benchmark

Put pressure on the path.
See what moves.

Switch on a workload, dial up the pressure, and watch a synthetic inference mix rebalance in real time. This is the behavior ACE is designed to make visible.

interactive benchmark / synthetic traffic

Stress the fabric.

Toggle the workload mix, push the input pressure up, then inject a burst — the fabric absorbs it while the unoptimized path flattens out.

simulated / live
01 modality mix
02 input pressure 68%
quietburst
inputs
Text01
Vision02
Audio03
Video04
fabric
outputs
served this window120 req/s
absorbed by fabric1,284 req
throughput / req per sec
773+40% vs. unoptimized path
with aceorigin path
p95 latencyms
171−18ms / adaptive route
token efficiency%
44context reused
cost / 1M tokensUSD
$0.151reclaimed by fabric
synthetic benchmark · 3 active modalities · steady state · sample 000
benchmark / 00synthetic traffic only
/ live arbitrage

The most cost-effective tier
capable of fulfilling the request.

Traffic fills committed capacity first and spills to metered only when it has to. These are measured fleet savings, not a projection.

merit-order arbitrage
incoming 10,000 QPSlive
01Azure OpenAI PTUPrimary Cloud APIcommitted throughput, saturated first$0.1585%
02OpenAI DirectElastic Burst APImetered spill once the PTU is full$0.5015%
03vLLM ClusteroptionalSelf-Hosted GPUsadd your own hardware and this rung fills first$0.04amortised—
live savings · projected
$5,274saved today

across 17.7M tokens

vs. no ladder
60%

$243.00/s vs $600.00/s

/ try it first-hand

Stop paying the token tax.
Reclaim your gross margins today.

Sign in to your developer console and generate your ACE_DEV_KEY on demand. No waitlist, no call, no credit card.

/ pick your first path

From request
to useful work.

Start where your stack is today. The fabric follows when your infrastructure grows up — same integration surface, more levers.

see the full onboarding doc
setup path / 0190% of users
Managed Cloud APIs

Drop in.
Keep moving.

Azure OpenAI, AWS Bedrock, GCP Vertex, OpenAI/Anthropic APIs.

Requestpii_ner · semantic_cache
Workloadllm_router
Stackstorm_guards
Paste API key + base URL live in 60 seconds · 4 skills
/ how the fabric thinks

Infrastructure that behaves
like a material.

Every layer of ACE is built around the same question: can this request do less work and still deliver more?

✳
behavior / route

Give every request the right path.

Adaptive routing weighs intent, context, latency, and model economics before a token is spent.

fabric_stateadaptivedecision_latency12mscontext_reuseenabled
/ the quiet advantage

Less motion.
More useful work.

ACE sits below your application and above the raw economics of inference. It finds the shortest useful path between intent and output, then keeps the system honest as demand changes.

01one integration surface
∞models, providers, workloads
00platform rewrites required
read the architecture note
/ live developer surface

Try the fabric
before you wire it in.

Point your existing OpenAI or Anthropic client at the ACE proxy URL. Caching, routing, pruning and guardrails activate automatically — your SDK calls, streaming and tool use stay identical.

ace / drop-in proxy wire-compatible
supportedOpenAI · Anthropic · Azure OpenAI SDKs supported
supportedChat Completions and the Responses API (GPT-5 reasoning)
supportedStreaming, tool-calling, thinking and vision preserved
supportedZero-config semantic caching via Redis / Qdrant
supportedPrometheus + OTLP telemetry out of the box
Semantic cache embeddings are computed locally via ONNX runtimes inside the proxy buffer. No prompt data leaves your perimeter for a cache lookup.
# OpenAI Python SDK — swap base_url, and use your ACE key
from openai import OpenAI

client = OpenAI(
    api_key="ace_dev_...",                      # not your OpenAI key
    base_url="https://engine.acefleet.dev/v1",  # 👈 ACE intercepts
)

resp = client.chat.completions.create(
    model="gpt-5",
    messages=[{"role": "user", "content": heavy_prompt}],
)
app.pyzero translation gateway
/ configure your skills

Turn the fabric
on gradually.

Every optimization skill ships with a three-tier trust ladder. Start in shadow — a dry run on real traffic with zero output mutation — inspect the behavior, then promote only the paths the numbers justify.

prodsemantic_cache
cost saving / semantic_cache

Vector similarity match at ≥0.93 threshold to replay responses at $0.

Serves cached responses instantly at 0 token cost.

3 live · 2 shadow
reversible observable wire-compatibleskill reference
/ what ace does

Cut API costs without
changing your architecture.

Three deployment surfaces, one control plane. 12 skills compose underneath.

target50–70% saved

API & Agentic Applications

AI app builders, agent developers, and early-stage startups consuming third-party frontier APIs.

prompt_tokens−58%
cache_hit_ratio0.71
avg_latency18ms
cost_per_1k_calls$0.34
01

Semantic Prompt Caching (sub-20ms)

Intercepts repetitive user queries and returns cached responses instantly at zero token cost.

02

Dynamic Intent-Based LLM Routing

Routes multi-step reasoning to GPT-4o / Claude 3.5 Sonnet; offloads JSON extraction, formatting, and classification to Llama 3.1 8B / GPT-4o-mini.

03

Entropy-Based Prompt Pruning (LLMLingua-2)

2×–5× input token reduction while preserving generation quality.

04

Agentic KV Summarizer & Loop Guardrails

Stops exponential context-length scaling during long agent execution loops.

observabilityThe receipt, not the promise.

See your exact dollar savings in real time.

ace / dashboard / overview
tokens_saved
184.2M
+23.4%· last 7 days
reclaimed_spend
$27,481
+31.7%· vs baseline
cache_hit_ratio
62.8%
+4.1pp· rolling 24h
latency_delta
−38ms
p50 faster· vs upstream
reclaimed_spend · usd / hourlast 18h
peak
$1,842 / hr
avg
$1,140 / hr
since deploy
$184,720
request_disposition
  • semantic_cache47%
  • llm_router31%
  • prompt_compaction16%
  • passthrough6%
cost_attribution · by_provider
anthropicnet $9,240 · saved $7,180
openainet $6,120 · saved $5,240
google_vertexnet $3,480 · saved $2,310
deepseeknet $1,260 · saved $940
cost_attribution · by_model_endpoint
gpt-5.6-solnet $6,120 · saved $5,240
fable-5net $5,180 · saved $4,320
gemini-3.6-flashnet $2,940 · saved $1,980
deepseek-v4-pronet $1,040 · saved $780
router_intelligence · why_it_routed● live
  • optimized_for_cost58%
    simple prompt → cheaper model
    n = 12,481
  • optimized_for_latency27%
    primary lagging → faster peer
    n = 5,812
  • fallback / auto_retry15%
    upstream 5xx → sibling provider
    n = 3,204
traffic_switch · gpt-5.6-sol ⇢ fable-5last 24m
openai 503 detected · auto-switched · 0ms downtime
gpt-5.6-sol (primary)fable-5 (failover)
fail_open_guarantee● active

ACE is designed to be invisible. If any internal skill exceeds its latency budget, or the gateway itself encounters an error, ACE instantly degrades to a direct pass-through — your request is relayed straight to the downstream provider, pipeline bypassed, upstream body untouched.

x-ace-served-by: direct_passthroughtrigger: skill > 50ms · panic · storage unreachableproduction traffic: never drops
request_lifecycle · most_recenthover a row
requesttypemodel_routed_tocachecostsaved
Hey, could you please refactor the auth module across all of the 12 files to use the new session APIcodingfable-5 → deepseek-v4-proMISS$0.3040 $0.0101$0.2939
Given these three conflicting lab results, can you tell me which hypothesis best explains the data?reasoninggpt-5.6-sol → kimi-k3MISS$0.0700 $0.0366$0.0334
I need you to prove that the sum of the first n odd integers equals n squaredformal_mathgpt-5.6-sol → deepseek-r1MISS$0.0975 $0.0073$0.0902
Quick question — what is the difference between gross margin and contribution margin?knowledge_qasemantic_cacheHIT$0.0235 $0.0000$0.0235
Review this 340-page vendor agreement very carefully and flag every auto-renewal clauselong_docopus-4.8 → gemini-3.6-flashMISS$0.6800 $0.2040$0.4760
Navigate to the billing portal for me and download last month's invoicecomputer_usefable-5 → sonnet-5MISS$0.0840 $0.0252$0.0588
Please just classify this support ticket by urgency and product areaclassifygemini-3.6-flash → deepseek-v4-proMISS$0.0006 $0.0002$0.0004
Hi, I wanted to ask — how do I rotate an expired API key without downtime?support_faqsemantic_cacheHIT$0.0280 $0.0000$0.0280
Sorry to bother, but what is the parental leave policy for US-based employees?policy_ragsemantic_cacheHIT$0.0370 $0.0000$0.0370
hover a request to trace its lifecycle · struck text is what compaction removed, struck cost is what the caller's model would have billed
try it liveOne prompt, the full engine trace.

LIVE DEMO.

The left pane streams the engine's decision log; the right pane returns the model response with token and cost accounting.

ace / console / try
free trial · 10 of 10 sends left today
Engine decision log
[req-a41f9c2d80b3] [in]: "implement fizzbuzz in typescript with tests"
[req-a41f9c2d80b3] [pii]: nothing to redact · ner_regex + ner_model scanned
[req-a41f9c2d80b3] [router]: query_category=code · complexity=medium → gpt-5-mini
[req-a41f9c2d80b3] [guard]: clean · pattern rules + classifier · classifier in shadow mode
[req-a41f9c2d80b3] [cache]: MISS · best-similarity=0.0000 (threshold 0.9200) · scope=session
[req-a41f9c2d80b3] [compact]: -0 tok · below the 40-token floor, sent unchanged
[req-a41f9c2d80b3] [serve]: gpt-5-mini @ azure-v1 · tok 10/42 (16 reasoning) · $0.000087
Responsegpt-5-mini
function fizzbuzz(n: number): string { if (n % 15 === 0) return 'FizzBuzz'; if (n % 3 === 0) return 'Fizz'; if (n % 5 === 0) return 'Buzz'; return String(n); }
tok in
10
tok out
42
cost
$0.000087
reasoning 16 tok
Router selectedgpt-5-mini
Unrealized savings$0.000870 / request
Cheaper than claude-opus-4-5$0.013914 / request
BillingBYOK_PASS_THROUGH

log lines and header values captured verbatim from the ACE production engine · model catalog 07312026

Ready to run this on your production traffic?
Start Free Trial →
watch the demo

See ACE route, cache, and save in real time.

A two-minute walkthrough of the production engine: one prompt, the full decision trace, and the cost accounting that shows up in your dashboard.

youtube · 0IOG8zqCd7w2 min walkthrough
ready when you are

Make every token
earn its way through.

Bring your real traffic. We will show you where the fabric gives work back.

open source/the local cost layer

Make every internal coding-agent token work harder.

ace-sidecar helps you cut what Claude Code costs you. It prices every turn on your own machine and ranks the changes worth making — no account, no upload.

  • Runs locally
  • Nothing leaves the machine
  • Drop-in proxy
$ 00 / what the sidecar did

Your spend, with the busywork removed.

Measured from the last 30 days of local agent traffic.

saved this month
$418.27

18.4% below raw spend

effective cost / turn
$0.034

from $0.042 without the sidecar

context reused
72.8%

1.2B tokens replayed locally

quality guardrail
98.6%

sessions within baseline

uv tool install ace-sidecar && ace up

Sample workspace from a reference project — the sidecar reports your own numbers, on your own machine. Python 3.12+ · AGPL-3.0 · loopback only · claude code, antigravity, codex. The sidecar measures one developer's sessions; ACE Fleet cuts the bill in the request path for a company's production traffic.

$01/how it works

One small layer between you and the bill.

Nothing changes in your CLI. ACE observes the request, applies safe optimizations, then forwards only what the model needs.

live path
  1. 01

    Keep using your CLI

    Claude Code, Codex, or Gemini CLI work exactly as before.

  2. 02

    ACE optimizes locally

    Cache hits, dedupe, and routing happen before a provider call.

    $418 saved
  3. 03

    Provider sees less

    Same intent, less repeated context, lower bill.

same algorithms as the gateway, open sourceThe cache, dedupe and routing here are the ACE gateway's own — the hosted path teams point production traffic at. ace-sidecar runs that same work at the other edge of the network: on your machine, against your coding agent, with the source open.

  • gateway hosted, production traffic
  • sidecar local, your coding agent
  • open source read it, fork it, run it

$418 is the 30-day figure from the reference project above, not live telemetry — what a repo saves depends on its own traffic.