← /blog
· ACE Engineering#compute-audit #mcp #headroom #finops #privacy #agents #managed-api-stack

Audit Your LLM Spend in One Prompt — Without Sending Us Any Code or Data

compute_audit is a free MCP tool with two parts: an in-house compute audit of what your workflow costs today, and an ACE gateway headroom analysis of what it could save. It runs in your environment; your workload never leaves it.

Most teams can tell you their monthly LLM bill. Few can tell you which line of code is driving it, or how much of it is waste.

compute_audit answers both. It is a free MCP tool that turns the coding agent you already run into an LLM cost auditor. Point it at one workflow and get a report in two parts.

Part 1 — In-house compute audit: what you spend today

A customised audit of your workflow as it runs right now:

  • What one task costs: calls, tokens by cache status, and spend by call class and billing line.
  • Where the waste is: prefixes rewritten mid-task, context the task never reads, expensive models on cheap jobs, output limits nobody tuned. Every finding comes with a file:line and a fix.
  • What you already do well: an efficient setup deserves credit, and knowing it tells you where not to spend effort.

Part 2 — ACE gateway headroom analysis: what you could save

For every ACE gateway skill (semantic cache, prompt compaction, trajectory folding, the model router, output and reasoning budgets), the analysis measures the redundancy in your workload and prices the ceiling. The combined saving is a share of spend, with every lever applied to what the earlier ones leave, so nothing is counted twice.

It ends with a skill plan your agent can apply directly. You read it as tables, and your agent reads it as JSON whose steps are ACE API calls.


Your workload never leaves your machine

The MCP server doesn't analyse anything. It sends your agent a method, and your agent runs it locally against your repo or logs.

  • Code, prompts, logs and results stay local. Nothing is uploaded.
  • Nothing is replayed through a gateway, and nothing is billed.
  • No account, no key. The public MCP server is read-only and can't place an inference call.
  • Every argument is optional. Call it with none, and even your workflow's name stays with you.

The one request is your agent fetching the playbook. Everything after that happens in your environment.

No logs? No problem. Without recorded requests, the audit reads the setup straight from your code: call sites, models, cache markers and loop shape. Anything that needs real traffic is marked not measured, never guessed.


Try it

claude mcp add --transport http ace https://acefleet.dev/mcp

Run compute_audit for the ticket-triage workflow on this codebase.


What you get

Illustrative example: a reference support-triage agent, not a customer workload.

Overall, 17–29% of ticket-triage spend can be saved. Estimated cost per task goes from 0.41to0.41 to0.29–$0.34.

Part 1 — found in the code

Finding Evidence Fix
A 312-entry knowledge-base index rides in every prompt prompts/system.py:41 Scope it per task
Search results appended whole, step after step tools/search.py:120 Fold stale history
Triage classifier on Sonnet; Haiku already tags the same tickets agent/classify.py:22 Route it

Part 2 — headroom per ACE skill

ACE skill Mode Saving
prompt_compaction prod 6–11%
agent_trajectory_compaction prod 8–15%
llm_router shadow → canary 4–6%
output_budget, semantic_cache shadow not measured
Combined 17–29%

The plan your agent applies (abridged):

{
  "schema": "ace.skill_plan/v1",
  "apply": [
    {
      "call": "POST /api/v1/dev_key/skills",
      "body": {
        "skill_modes": {
          "prompt_compaction": "prod",
          "agent_trajectory_compaction": "prod",
          "llm_router": "shadow"
        },
        "skill_params": { "agent_trajectory_compaction": { "expected_task_steps": 8 } }
      }
    }
  ],
  "savings": { "day_one": "14–24%", "after_rollout": "17–29%" }
}

New to ACE? Your agent starts at step one: store a provider key, mint a developer key, apply the plan. Already on ACE? It compares the plan with your current settings and applies only what differs.


One prompt. Two answers: what you spend, and what you could save. Zero bytes of your workload sent to us. Run it on your most expensive workflow today.