all docs

Prompt pruning

Lossless prompt compaction ahead of outbound model calls.

problem it solves
Eliminates format overhead and repeated lines that inflate input token bills: pretty-printed JSON, empty fields, ServiceNow link triples, whitespace and markdown decoration.

What it does

What it does: Rewrites every span of the prompt losslessly: JSON re-emitted compact (number literals and key order kept), empty fields dropped and ServiceNow `{display_value, value, link}` triples collapsed in JSON of 200+ characters, exact repeated lines dropped, whitespace and markdown decoration removed. No word is deleted.

What it covers: The system prompt, user turns and tool results — including Bedrock `toolResult.json` and Gemini `functionResponse` objects, which stay objects.

Under a prompt-cache breakpoint: Conversation spans are rewritten on the step they are live and replayed byte-identical on every later step, so the provider's cache entry holds. The system prompt, which other sessions share, is rewritten only when the cache arithmetic pays back (`x-ace-compaction-prefix`).

Word pruning (opt-in): `skill_params.prompt_compaction.prune_words: true` adds a keep-score pruner after the lossless pass. It deletes low-information words; instructions, negations, quantifiers, modal words and identifiers are protected.

How it recovers: A span the skill cannot finish inside its time ceiling goes as received and is pinned, so later steps send the same bytes.

What we need from you

  • Nothing to configurerequired

    The lossless pass runs in the gateway: no compactor model, no provider key, no per-model setup. Toggle it on for the key.

  • Sufficient prompt lengthrecommended

    Savings concentrate in tool output and pasted payloads; short prompts are left alone.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Prompts are transmitted to upstream models unmodified.No prompt_compaction stage recorded.
shadowComputes the rewrite and its saving without modifying the outbound request.Stage with action=would_compact; `x-ace-compaction-tokens-would-save`.
prodSends the rewritten prompt to the upstream model.`x-ace-compaction-tokens-saved` and `x-ace-compaction-spans` say how much and where.

Current policy

RewriteLossless only`prune_words: false`. Deletes no word.
Cached historyRewritten`protect_cached_history: false`. Set true when a key's sessions also go direct to the vendor.
Latency~5 ms per 10k prose tokensUnder 1 ms per 10k tokens of JSON; history is replayed from memory.

Worth knowing before you enable it

  • ·Rewriting cached history assumes every step of a session passes through ACE. A key that also sends sessions direct to the vendor should set `protect_cached_history: true`, or the direct route's cache sees history it never wrote.
  • ·A whole-JSON tool result is shown to the model compact. An agent that reads a JSON file and then edits it by exact string match would match against text the file does not contain.
  • ·`x-ace-body-rewritten: history` means conversation items up to the breakpoint changed and tools and system did not; `prefix` means a tool or system item changed.
  • ·On a semantic-cache hit or a completion replay the skill never runs — there is no upstream prompt to shrink — and the response says so: `x-ace-compaction: skipped` + `x-ace-compaction-reason: cache_hit` (or `replay_hit`) + `-mode`. Absent `x-ace-compaction-*` headers mean the skill is OFF for the key, never that a hit happened to answer first; read the two apart when judging whether compaction runs on a route.

What it replaces

  • ·Manual text summarization passes before API calls.
  • ·Custom regex token stripping scripts.
  • ·SDK-level conversation truncation helpers.

Custom pruning rules and domain models are available on the enterprise tier.

  • ·Domain-specific compaction models (code, medical, legal).
  • ·Custom element protection rules.
  • ·Adjustable compression aggressive sliders.
team@acefleet.dev →