Pay for savings.
Never for the plumbing.
Start free with a developer key and run real traffic through the production engine. Paid tiers are priced against the compute dollars ACE reclaims — if it saves nothing, you owe nothing.
When request limits are exceeded on any tier, skills automatically turn into shadow mode (quiet passthrough — zero actions taken on your requests). Eligible to re-activate in prod once limit refreshes or tier is upgraded.
A frictionless, top-of-funnel magnet — prove the value in your own terminal.
10,000–50,000 requests / logs per month
- ▸Base-URL swap, all providers
- ▸Standard context compaction
- ▸Lightweight semantic caching
- ▸Dynamic intent routing
- ▸Live cost-savings proof of concept
Basic caching + routing immediately demonstrates ROI on repetitive queries — a live preview of the 30–80% savings unlocked by heavier pruning on paid tiers.
The core self-serve revenue engine for established apps and tech leads.
1,000,000 standard requests + 100k GPU-pruned requests
* $0.50 per additional 1,000 pruned requests
$10.00 per additional 1,000,000 standard requests (or $0.01 per 1k)
- ▸Live telemetry dashboards
- ▸Full semantic caching + DB logging
- ▸Dynamic intent routing
- ▸Heavy context pruning (capped at 250k)
- ▸Transparent usage-based overage
Delivers the core promise: slash monthly API and GPU costs by 30–80% without sacrificing accuracy or speed — instantly reclaiming gross margins.
Massive scale and strict compliance — data never leaves your cloud.
Massive enterprise volume
- ▸VPC / on-premise deployment
- ▸Advanced governance (RBAC, SSO, custom PII masking)
- ▸Custom SLAs
- ▸GPU fleet merit-order dispatch
- ▸Named engineer, shared Slack
Applies the 30–80% compute reduction at enterprise scale — shaving token bloat translates to massive absolute dollar savings while keeping all data secure.
Verified savings = provider list price minus net invoice, measured on your own traffic in the control plane.