all docs

Agent trajectory compaction

Folds old agent turns into a checkpoint on a fixed grid, so a long tool loop stops re-sending its whole history — and the bytes it does send stay cache-stable between folds.

problem it solves
Stops exponential context growth in multi-turn agent tool loops from inflating token costs, without moving the prefix your provider has cached on every turn.

What it does

What it does: Partitions the conversation into turns — an assistant action plus everything that answers it — and seals the leading ones into a single `[trajectory summary]` block, one line per sealed turn, keeping the recent turns verbatim. No external LLM call — the summary is extractive, and every line names the tool-call id it stands for (`tool_result[toolu_3] query_table: …`) so an agent that re-reads results by id still can. A sealed tool result is carried as its id, tool name and a bounded head — 320 characters (`ACE_AGENT_TOOL_RESULT_HEAD_CHARS`, deployment-level; 0 = unbounded) — never the whole result, whatever the verbatim threshold says.

The fold is a pure function of the transcript. The fold point sits on a fixed grid — `sealed = E·k + 1 − R`, with `E` the epoch size (5) and `R` the turns kept verbatim (2) — so it moves at turns 6, 11, 16, … and nowhere else — and at each boundary, or when the body crosses the token watermark, only if it pays: the gate prices what the fold would rewrite (a cache write, 1.25×) against what it saves on every later step inside the key's payback horizon, and holds (`not_worth_it`) otherwise. Between two boundaries the summary text is byte-identical and your new rounds are appended after it, so the provider reads the prefix from cache; at a boundary the prefix is rewritten once. Two gateway replicas, or one before and after a restart, compute the same bytes from the same history without sharing anything.

Applied on every surface. The fold reaches the provider on `/v1/chat/completions`, `/anthropic/v1/messages`, `/gemini`, `/bedrock` and the Responses shim alike. The summary is inserted as a user turn and merged into an adjacent user turn where there is one, because Bedrock rejects two consecutive user turns and Gemini does on some paths.

What is never summarised — recognised from the shape of every request, so a client that rewrites its own history is simply re-read as it now is:

  • ·A system message, anywhere.
  • ·The opening user message — the task, in the caller's words, not the first sentence of them.
  • ·On a trajectory whose observations arrive as tool results (`tool_result` blocks, `role: tool`, `toolResult`, `functionResponse`), every user message that is not one: an injected reminder, a memory checkpoint the client built, a correction.
  • ·A text block riding beside tool results in the same message — the caller's instruction, not an observation.
  • ·Any message at or under 240 characters, carried into the summary verbatim rather than cut to a sentence: a fold pointer, a one-line status, a marker.
  • ·A turn carrying a cache breakpoint, kept whole — summarising the results that answer a marked assistant message would orphan a `tool_use`.
  • ·The live turn and the one before it, always the caller's own objects byte for byte, marker and signed thinking block included.

It compacts every turn once it starts, and your client never sees it. Your app keeps its own full history and keeps sending it; the gateway recomputes the fold on each request and sends the shortened version upstream.

Every response says what it did. `x-ace-trajectory-compaction` carries the verdict (`compacted`, `would_compact`, `no_op`, `skipped`, `refused`, `failed`) with `-reason`, `-strategy`, `-turns` (`sealed/total`), the message and token counts, and `-checkpoint` — a 12-hex digest of the summary text that changes exactly when the folded prefix changes, so you can tell a fold-induced cache write from any other kind against the provider's `cache_creation_input_tokens`.

What we need from you

  • Multi-turn agent sessionsrequired

    Requires agent interaction patterns using tool calls and observation cycles.

  • A session handle (recommended)recommended

    Send `x-ace-session: <task id>` (or `session_id` in the body) on every call of a task. The id is yours to choose — any stable id for the task or agent run. It may span several transcripts: when your app rebuilds its history (a fresh memory summary as the opening message) or runs sub-agents side by side, the fold keeps its state per transcript inside the session. Without a handle the gateway derives one from the transcript's opening messages, which every call of one transcript re-sends, so a caller who sends none is still folded; the handle is what groups the task's calls in telemetry and the knowledge graph.

  • One handle per task, not per userrecommended

    `user` is accepted last because it names a person, not a task; two tasks of one user must not share a handle.

  • Prompt pruning skill enabledrecommended

    Pairs with prompt pruning for maximum context reduction.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Full trajectory history is sent on every step.No agent_trajectory_compaction stage recorded; no x-ace-trajectory-compaction header.
shadowComputes the fold and measures it without altering the payload.`x-ace-trajectory-compaction: would_compact` with `-tokens-would-save` (and `-tokens-saved: 0`); the trace stage carries action=would_compact, sealed and total turns, and the checkpoint digest.
prodSends the folded conversation upstream, on every surface.`x-ace-trajectory-compaction: compacted` with `-turns`, `-messages-before/-after`, `-tokens-before/-after/-saved`, `-checkpoint`, `-write-tokens` (what the fold cost to rewrite, on the step it moved) and `-held` (fold points the gate refused so far). The folded conversation is read back per request with `GET /api/v1/tenant/request/{request_id}?messages=true` (`messages_compacted`, tool calls and results as parts); list views — the durable `/api/v1/tenant/requests` and the in-memory ring — carry only `messages_compacted_count`.

Current policy

Epoch size5 turns (`epoch_size`, floor 2)The fold point moves at turns 6, 11, 16, … and nowhere else. The first turns of a session are a `no_op` (`first_epoch`) by design.
Kept verbatimThe live turn plus the one before it (`keep_recent_turns`, floor 1)Images in older turns go with their turn when it is sealed (`keep_recent_images`, 1).
Early fold20,000 tokens (`token_watermark`; 0 = off)A trajectory crossing this is offered to the gate before the grid would have, and is priced like any boundary — a watermark reached past `expected_task_steps` never fires, because at one step of horizon nothing pays. The latch is the only per-session state, it only moves the fold point forward, and losing it costs one re-fold back to the grid — never a different summary of the same turns. Set 0 for strict restart-stability.
Payback horizon`expected_task_steps` 12 (floor 1); `payback_turns` unset (floor 1)How many later reads a fold may count on. `expected_task_steps` caps the horizon at what is left of the task, so at the default 12 the first grid boundary is taken (six steps left) and every fold point after step 12 has one step to pay back and is held. A key whose tasks run 30–80 steps sets it to the length it measures (`40`): every boundary is then priced with the steps that remain and taken while it pays. `payback_turns` pins the horizon outright ("pays back in 5 steps or not"), wherever in the task the gate is asked.
Verbatim threshold240 charactersA message at or under this is carried into the summary as-is rather than cut to a sentence.
Loop guardOff (`max_tool_repeats: 0`)A key that sets it and trips it is served unfolded with `refused: tool_loop` — reported, not swallowed. Halting a real loop is the multi-agent guard's job; an agent re-polling a status endpoint three times is not looping.
Execution engineDeterministic extractive parserSub-0.3ms execution overhead (<0.3ms p99), zero LLM model calls on critical path.

Worth knowing before you enable it

  • ·There is no weaker anonymous mode any more. A stock SDK caller with no session handle gets the same grid fold a named session gets — the fold is a function of the transcript, so there is no shared state for two anonymous conversations to contaminate. `-strategy` reads `epoch`; `sliding_window` appears only for a compactor deployed in that mode.
  • ·It does not shrink YOUR context window. The compaction applies to what we send the provider; your app still holds and sends its full history, so it hits its own limit on the same schedule as before.
  • ·Extended thinking no longer refuses the skill. The fold never rewrites a turn — it drops whole sealed turns and keeps recent ones by reference — and the provider verifies only the last assistant turn's signature, which is always retained. Prompt compaction still refuses on a signed block, because it rewrites prose.
  • ·On a trajectory whose observations arrive as plain user messages (not tool results) the fold cannot tell an observation from an injected instruction, and summarises both.
  • ·Once a conversation has folded, every later turn in it is served from the checkpoint. That is the design, not a bug: it keeps the bytes we send byte-identical between folds, which is what keeps a cached prefix alive.
  • ·How aggressively to fold is these parameters, not a mode: `expected_task_steps` (the one that matters — set it to the task length you measure), `payback_turns` (a fixed horizon), `epoch_size` (how often the gate is asked; smaller means more, smaller folds, each moving the prefix once), `token_watermark` (the off-grid fold for large, few steps; 0 = off, and the fold is then a pure function of the transcript) and `keep_recent_turns` (the retained window). Cautious is the defaults; aggressive is `expected_task_steps` at the measured length, `epoch_size` 5, `token_watermark` 20000, `keep_recent_turns` 1–2.
  • ·Per-key params (`epoch_size`, `keep_recent_turns`, `keep_recent_images`, `token_watermark`, `max_tool_repeats`, `expected_task_steps`, `payback_turns`) are set through `POST /api/v1/dev_key/skills` with body `skill_params: {agent_trajectory_compaction: {…}}`, alongside `skills` and `skill_modes`, and are validated against floors (`epoch_size` ≥ 2, `keep_recent_turns` ≥ 1, `expected_task_steps` and `payback_turns` ≥ 1, the rest ≥ 0; a value below its floor is a 400 naming the field). `skill_params` is Enterprise-only: on any other plan the whole save is a 403, `skill_params is only configurable on the Enterprise plan`. The defaults above are what every key runs without them. `ACE_AGENT_MAX_TURNS` is no longer read.
  • ·Extremely long observation outputs (like full raw HTML pages) should be pre-filtered.

What it replaces

  • ·Hand-written agent memory windowing logic.
  • ·Custom loop detection and cancellation counters.
  • ·Manual history truncation in LangChain / LlamaIndex workflows.

Advanced agent guardrails and state persistence are available on enterprise.

  • ·Custom agent state schema definitions.
  • ·Cross-session persistent agent memory stores.
  • ·Configurable loop-detection heuristics.
  • ·Per-key policy through `skill_params.agent_trajectory_compaction` on `POST /api/v1/dev_key/skills` — the epoch size, what is kept verbatim, the payback horizon, the early-fold watermark and the loop guard, each above its floor.
team@acefleet.dev →