Pay for savings.
Never for the plumbing.
Start free with a developer key and run real traffic through the production engine.
Over your limit, skills drop to shadow mode — requests pass straight through, untouched. They resume when the limit resets or you upgrade.
Try it on real traffic.
10k–50k requests / month
Caching + routing — a preview of the full 30–80%
For production apps and teams.
1M standard + 100k GPU-pruned requests / month
$0.50 per 1k pruned requests
$10 per 1M standard requests
no hard caps — usage-based
30–80% off API + GPU spend, same accuracy and speed
Your cloud. Data never leaves it.
Unlimited
- ▸VPC / on-premise deployment
- ▸RBAC, SSO, custom PII masking
- ▸Custom SLAs
- ▸GPU fleet merit-order dispatch
- ▸Named engineer, shared Slack
The same 30–80%, at enterprise scale
Verified savings = provider list price minus net invoice, measured on your own traffic in the control plane.
How much committed capacity
should you actually buy?
Azure PTUs, Bedrock Model Units and Vertex GSUs are cheaper per token and bill around the clock — so the question is never "how much do we use" but "how much of the day is worth pre-paying for." Answer four questions and find out. No key, no account, nothing leaves your browser.
This sizes the commitment you make to your provider. It is not a quote for ACE.
What you pay OpenAI, Azure, Bedrock or Vertex today, at list.
Implied blended rate $0.0067/1k.
Peak-to-trough ratio. Without it there is no distribution to size against, and the honest answer to every question becomes "commit everything."
Your quoted rate against list. This moves the answer more than anything else on this page — the optimal service level is exactly 1 − committed/metered.
Committed throughput to buy (4 machines / PTU units at $2,000/mo floor). Run it at 5,621 tok/s — the extra 20% is headroom, because past 80% utilization latency jitter breaks the SLO and the router starts steering traffic away.
estimate — estimated from the figures you entered — not measured against your traffic. The same newsvendor sizing runs against your real per-hour demand once traffic flows through the gateway, and that answer supersedes this one.
Sign Up for a Free Compute Audit →Get sizing verified against your real traffic · no credit card