Audit Your LLM Spend in One Prompt — Without Sending Us Any Code or Data
compute_audit is a free MCP tool with two parts: an in-house compute audit of what your workflow costs today, and an ACE gateway headroom analysis of what it could save. It runs in your environment; your workload never leaves it.
Most teams can tell you their monthly LLM bill. Few can tell you which line of code is driving it, or how much of it is waste.
compute_audit answers both. It is a free MCP tool that turns the coding agent you already run
into an LLM cost auditor. Point it at one workflow and get a report in two parts.
Part 1 — In-house compute audit: what you spend today
A customised audit of your workflow as it runs right now:
- What one task costs: calls, tokens by cache status, and spend by call class and billing line.
- Where the waste is: prefixes rewritten mid-task, context the task never reads, expensive
models on cheap jobs, output limits nobody tuned. Every finding comes with a
file:lineand a fix. - What you already do well: an efficient setup deserves credit, and knowing it tells you where not to spend effort.
Part 2 — ACE gateway headroom analysis: what you could save
For every ACE gateway skill (semantic cache, prompt compaction, trajectory folding, the model router, output and reasoning budgets), the analysis measures the redundancy in your workload and prices the ceiling. The combined saving is a share of spend, with every lever applied to what the earlier ones leave, so nothing is counted twice.
It ends with a skill plan your agent can apply directly. You read it as tables, and your agent reads it as JSON whose steps are ACE API calls.
Your workload never leaves your machine
The MCP server doesn't analyse anything. It sends your agent a method, and your agent runs it locally against your repo or logs.
- Code, prompts, logs and results stay local. Nothing is uploaded.
- Nothing is replayed through a gateway, and nothing is billed.
- No account, no key. The public MCP server is read-only and can't place an inference call.
- Every argument is optional. Call it with none, and even your workflow's name stays with you.
The one request is your agent fetching the playbook. Everything after that happens in your environment.
No logs? No problem. Without recorded requests, the audit reads the setup straight from your
code: call sites, models, cache markers and loop shape. Anything that needs real traffic is
marked not measured, never guessed.
Try it
claude mcp add --transport http ace https://acefleet.dev/mcp
Run
compute_auditfor theticket-triageworkflow on this codebase.
What you get
Illustrative example: a reference support-triage agent, not a customer workload.
Overall, 17–29% of ticket-triage spend can be saved. Estimated cost per task goes from
0.29–$0.34.
Part 1 — found in the code
| Finding | Evidence | Fix |
|---|---|---|
| A 312-entry knowledge-base index rides in every prompt | prompts/system.py:41 |
Scope it per task |
| Search results appended whole, step after step | tools/search.py:120 |
Fold stale history |
| Triage classifier on Sonnet; Haiku already tags the same tickets | agent/classify.py:22 |
Route it |
Part 2 — headroom per ACE skill
| ACE skill | Mode | Saving |
|---|---|---|
prompt_compaction |
prod | 6–11% |
agent_trajectory_compaction |
prod | 8–15% |
llm_router |
shadow → canary | 4–6% |
output_budget, semantic_cache |
shadow | not measured |
| Combined | 17–29% |
The plan your agent applies (abridged):
{
"schema": "ace.skill_plan/v1",
"apply": [
{
"call": "POST /api/v1/dev_key/skills",
"body": {
"skill_modes": {
"prompt_compaction": "prod",
"agent_trajectory_compaction": "prod",
"llm_router": "shadow"
},
"skill_params": { "agent_trajectory_compaction": { "expected_task_steps": 8 } }
}
}
],
"savings": { "day_one": "14–24%", "after_rollout": "17–29%" }
}
New to ACE? Your agent starts at step one: store a provider key, mint a developer key, apply the plan. Already on ACE? It compares the plan with your current settings and applies only what differs.
One prompt. Two answers: what you spend, and what you could save. Zero bytes of your workload sent to us. Run it on your most expensive workflow today.