all docs

Output budget

A per-key ceiling on the completion budget a request declares. Only lowers; shadow reports.

problem it solves
Output tokens bill at 4–8× input, and a runaway generation is billed to whatever max_tokens the client sent — 32k on every step of a loop whose steps answer in 400.

What it does

What it does: Caps `max_tokens` / `max_completion_tokens` / `max_output_tokens` / `maxOutputTokens` at the key's ceiling when a request asks for more, in the request's own field, on every provider surface.

Never up: A request under the ceiling is sent as received. A request that declared no budget is left alone unless `default_max_output_tokens` is set.

A floor under the ceiling: Never caps below `min_output_tokens`, so a mis-set ceiling cannot make every answer a truncated one.

Says what it did: `x-ace-output-budget: capped | set | no_op | would_cap` with `-requested` and `-applied`.

What we need from you

  • A completion budget field on the surfacerequired

    `max_tokens` / `max_completion_tokens` (OpenAI, Anthropic), `max_output_tokens` (Responses), `inferenceConfig.maxTokens` (Bedrock), `maxOutputTokens` (Gemini). Every relayed surface has one.

  • A declared completion budgetrecommended

    Acts on the budget the request carries; `default_max_output_tokens` gives one to a request that carries none.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. The request's budget goes upstream as sent.No output_budget stage recorded.
shadowComputes the cap and sends the request as received.Stage with action=would_cap, before and after; `x-ace-output-budget: would_cap`.
prodRewrites the budget field to the ceiling when the request is above it.Stage with action=capped; `x-ace-output-budget-applied` is what went upstream.

Current policy

Ceiling16,384 tokens (`skill_params.output_budget.max_output_tokens`)Clears a large file write from a coding agent; bounds a runaway at half the usual 32k ask.
Floor1,024 tokens (`min_output_tokens`)The cap never goes below this.
Default for a request with noneoff (`default_max_output_tokens: 0`)Giving a budget to a caller who declared none changes a request that was correct as sent.

Worth knowing before you enable it

  • ·A capped request that then stops at the cap comes back with `finish_reason: length` / `stop_reason: max_tokens`, as it would have at the client's own cap. Set the ceiling above the longest answer the workload legitimately needs.
  • ·Anthropic requires `max_tokens`, so every request on that surface declares one; OpenAI does not, so `default_max_output_tokens` is the only way to bound a request that sends none.

What it replaces

  • ·Per-client max_tokens conventions nobody enforces.

Per-use-case ceilings and a truncation audit.

  • ·A ceiling per `x-ace-use-case` rather than per key.
  • ·A report of every capped request that stopped at the cap.
team@acefleet.dev →