Output budget
A per-key ceiling on the completion budget a request declares. Only lowers; shadow reports.
problem it solves
Output tokens bill at 4–8× input, and a runaway generation is billed to whatever max_tokens the client sent — 32k on every step of a loop whose steps answer in 400.
What it does
What it does: Caps `max_tokens` / `max_completion_tokens` / `max_output_tokens` / `maxOutputTokens` at the key's ceiling when a request asks for more, in the request's own field, on every provider surface.
Never up: A request under the ceiling is sent as received. A request that declared no budget is left alone unless `default_max_output_tokens` is set.
A floor under the ceiling: Never caps below `min_output_tokens`, so a mis-set ceiling cannot make every answer a truncated one.
Says what it did: `x-ace-output-budget: capped | set | no_op | would_cap` with `-requested` and `-applied`.
What we need from you
- A completion budget field on the surfacerequired
`max_tokens` / `max_completion_tokens` (OpenAI, Anthropic), `max_output_tokens` (Responses), `inferenceConfig.maxTokens` (Bedrock), `maxOutputTokens` (Gemini). Every relayed surface has one.
- A declared completion budgetrecommended
Acts on the budget the request carries; `default_max_output_tokens` gives one to a request that carries none.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. The request's budget goes upstream as sent. | No output_budget stage recorded. |
| shadow | Computes the cap and sends the request as received. | Stage with action=would_cap, before and after; `x-ace-output-budget: would_cap`. |
| prod | Rewrites the budget field to the ceiling when the request is above it. | Stage with action=capped; `x-ace-output-budget-applied` is what went upstream. |
Current policy
| Ceiling | 16,384 tokens (`skill_params.output_budget.max_output_tokens`) | Clears a large file write from a coding agent; bounds a runaway at half the usual 32k ask. |
| Floor | 1,024 tokens (`min_output_tokens`) | The cap never goes below this. |
| Default for a request with none | off (`default_max_output_tokens: 0`) | Giving a budget to a caller who declared none changes a request that was correct as sent. |
Worth knowing before you enable it
- ·A capped request that then stops at the cap comes back with `finish_reason: length` / `stop_reason: max_tokens`, as it would have at the client's own cap. Set the ceiling above the longest answer the workload legitimately needs.
- ·Anthropic requires `max_tokens`, so every request on that surface declares one; OpenAI does not, so `default_max_output_tokens` is the only way to bound a request that sends none.
What it replaces
- ·Per-client max_tokens conventions nobody enforces.
Per-use-case ceilings and a truncation audit.
- ·A ceiling per `x-ace-use-case` rather than per key.
- ·A report of every capped request that stopped at the cap.