API & Agentic Applications
AI app builders, agent developers, and early-stage startups consuming third-party frontier APIs.
Route less. Reuse more. Give every request the right model, context, and compute.
60s Setup (OpenAI, Azure, Anthropic). One integration surface, no application rewrite.
→ Every prompt cached, routed, and pruned before it burns a token.
Change one line. Your SDK calls, streaming and tool use stay exactly as they are.
base_url = "https://engine.acefleet.dev/v1"Every layer of your AI stack — from silicon to agent frameworks — adds cost, latency, and fragility. ACE operates across all of them.
Drop in your monthly spend and see what ACE would have saved on it. Cascaded waterfall math — no overlapping percentages, no black-box multipliers.
of metered spend, modelled on 30% cacheable · 30% routable · 30% compressible traffic.
Switch on a workload, dial up the pressure, and watch a synthetic inference mix rebalance in real time. This is the behavior ACE is designed to make visible.
Toggle the workload mix, push the input pressure up, then inject a burst — the fabric absorbs it while the unoptimized path flattens out.
Traffic fills committed capacity first and spills to metered only when it has to. These are measured fleet savings, not a projection.
across 17.7M tokens
$243.00/s vs $600.00/s
Sign in to your developer console and generate your ACE_DEV_KEY on demand. No waitlist, no call, no credit card.
Start where your stack is today. The fabric follows when your infrastructure grows up — same integration surface, more levers.
Azure OpenAI, AWS Bedrock, GCP Vertex, OpenAI/Anthropic APIs.
Every layer of ACE is built around the same question: can this request do less work and still deliver more?
Adaptive routing weighs intent, context, latency, and model economics before a token is spent.
ACE sits below your application and above the raw economics of inference. It finds the shortest useful path between intent and output, then keeps the system honest as demand changes.
Point your existing OpenAI or Anthropic client at the ACE proxy URL. Caching, routing, pruning and guardrails activate automatically — your SDK calls, streaming and tool use stay identical.
# OpenAI Python SDK — swap base_url, and use your ACE key
from openai import OpenAI
client = OpenAI(
api_key="ace_dev_...", # not your OpenAI key
base_url="https://engine.acefleet.dev/v1", # 👈 ACE intercepts
)
resp = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": heavy_prompt}],
)Every optimization skill ships with a three-tier trust ladder. Start in shadow — a dry run on real traffic with zero output mutation — inspect the behavior, then promote only the paths the numbers justify.
Serves cached responses instantly at 0 token cost.
Three deployment surfaces, one control plane. 12 skills compose underneath.
AI app builders, agent developers, and early-stage startups consuming third-party frontier APIs.
Intercepts repetitive user queries and returns cached responses instantly at zero token cost.
Routes multi-step reasoning to GPT-4o / Claude 3.5 Sonnet; offloads JSON extraction, formatting, and classification to Llama 3.1 8B / GPT-4o-mini.
2×–5× input token reduction while preserving generation quality.
Stops exponential context-length scaling during long agent execution loops.
ACE is designed to be invisible. If any internal skill exceeds its latency budget, or the gateway itself encounters an error, ACE instantly degrades to a direct pass-through — your request is relayed straight to the downstream provider, pipeline bypassed, upstream body untouched.
| request | type | model_routed_to | cache | cost | saved |
|---|---|---|---|---|---|
| Hey, could you please refactor the auth module across all of the 12 files to use the new session API | coding | fable-5 → deepseek-v4-pro | MISS | $0.3040 $0.0101 | $0.2939 |
| Given these three conflicting lab results, can you tell me which hypothesis best explains the data? | reasoning | gpt-5.6-sol → kimi-k3 | MISS | $0.0700 $0.0366 | $0.0334 |
| I need you to prove that the sum of the first n odd integers equals n squared | formal_math | gpt-5.6-sol → deepseek-r1 | MISS | $0.0975 $0.0073 | $0.0902 |
| Quick question — what is the difference between gross margin and contribution margin? | knowledge_qa | semantic_cache | HIT | $0.0235 $0.0000 | $0.0235 |
| Review this 340-page vendor agreement very carefully and flag every auto-renewal clause | long_doc | opus-4.8 → gemini-3.6-flash | MISS | $0.6800 $0.2040 | $0.4760 |
| Navigate to the billing portal for me and download last month's invoice | computer_use | fable-5 → sonnet-5 | MISS | $0.0840 $0.0252 | $0.0588 |
| Please just classify this support ticket by urgency and product area | classify | gemini-3.6-flash → deepseek-v4-pro | MISS | $0.0006 $0.0002 | $0.0004 |
| Hi, I wanted to ask — how do I rotate an expired API key without downtime? | support_faq | semantic_cache | HIT | $0.0280 $0.0000 | $0.0280 |
| Sorry to bother, but what is the parental leave policy for US-based employees? | policy_rag | semantic_cache | HIT | $0.0370 $0.0000 | $0.0370 |
The left pane streams the engine's decision log; the right pane returns the model response with token and cost accounting.
log lines and header values captured verbatim from the ACE production engine · model catalog 07312026
A two-minute walkthrough of the production engine: one prompt, the full decision trace, and the cost accounting that shows up in your dashboard.
Bring your real traffic. We will show you where the fabric gives work back.
Introducing Skill Knowledge Graph: an architectural breakthrough that transforms ephemeral multi-turn agent exploration into a durable, self-improving knowledge graph—slashing trial-and-error compute waste while elevating agent success rates.
How our automated model market catalog has been engineered to be real-time, air-gapped, and consolidated: delivering instant model onboarding, exposure across 216+ models in the LLM Router dropdown, strict FinOps counterfactual auditability, and zero-socket offline execution.
Architecting a high-throughput Bring-Your-Own-Key (BYOK) vault for multi-tenant AI gateways. How dual-tier AES-GCM-256 cryptographic context caching delivers sub-millisecond key resolution without exposing plaintext credentials.
open source/the local cost layer
ace-sidecar helps you cut what Claude Code costs you. It prices every turn on your own machine and ranks the changes worth making — no account, no upload.
Measured from the last 30 days of local agent traffic.
18.4% below raw spend
from $0.042 without the sidecar
1.2B tokens replayed locally
sessions within baseline
uv tool install ace-sidecar && ace upSample workspace from a reference project — the sidecar reports your own numbers, on your own machine. Python 3.12+ · AGPL-3.0 · loopback only · claude code, antigravity, codex. The sidecar measures one developer's sessions; ACE Fleet cuts the bill in the request path for a company's production traffic.
$01/how it works
Nothing changes in your CLI. ACE observes the request, applies safe optimizations, then forwards only what the model needs.
Claude Code, Codex, or Gemini CLI work exactly as before.
Cache hits, dedupe, and routing happen before a provider call.
Same intent, less repeated context, lower bill.
same algorithms as the gateway, open sourceThe cache, dedupe and routing here are the ACE gateway's own — the hosted path teams point production traffic at. ace-sidecar runs that same work at the other edge of the network: on your machine, against your coding agent, with the source open.
$418 is the 30-day figure from the reference project above, not live telemetry — what a repo saves depends on its own traffic.