all docs

Reasoning effort

A per-key ceiling on reasoning: a level on OpenAI, a thinking budget on Anthropic, Bedrock and Gemini. Never on, never up.

problem it solves
Reasoning tokens are output tokens, billed at 4–8× input, and their volume scales with whatever effort or budget the client happened to send on every step.

What it does

What it does: Lowers `reasoning_effort` / `reasoning.effort` to the key's ceiling (default `medium`) and caps `thinking.budget_tokens` / `thinkingBudget` at the key's ceiling (default 8,192) when a request asks for more.

Never on, never up: A request that declared no reasoning, or disabled it, is sent as received. A request under the ceiling is sent as received.

A floor under the ceiling: A thinking budget is never capped below `min_budget_tokens` (1,024, the least Anthropic accepts).

Says what it did: `x-ace-reasoning-effort: lowered | no_op | would_lower` with `-requested` and `-applied`.

What we need from you

  • A reasoning control on the requestrequired

    OpenAI `reasoning_effort` or `reasoning.effort`; Anthropic / Bedrock `thinking.budget_tokens`; Gemini `thinkingConfig.thinkingBudget`.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. The request's reasoning control goes upstream as sent.No reasoning_effort stage recorded.
shadowComputes the ceiling and sends the request as received.Stage with action=would_lower, before and after; `x-ace-reasoning-effort: would_lower`.
prodRewrites the level or the budget to the ceiling when the request is above it.Stage with action=lowered; `x-ace-reasoning-effort-applied` is what went upstream.

Current policy

Level ceilingmedium (`skill_params.reasoning_effort.max_effort`)What the GPT-5 family runs at when nothing is said; `minimal` < `low` < `medium` < `high`.
Budget ceiling8,192 tokens (`max_budget_tokens`)A generous thinking budget for most tasks.
Budget floor1,024 tokens (`min_budget_tokens`)The least Anthropic accepts; the cap never goes below it.

Worth knowing before you enable it

  • ·A lowered thinking budget changes the answer on a task that needed the reasoning. Start in shadow and read `-requested` against the task outcomes before moving to prod.
  • ·Anthropic requires `max_tokens` to exceed `thinking.budget_tokens`; lowering the budget never violates that, raising the ceiling above the request's `max_tokens` is refused by the provider, not by this skill.

What it replaces

  • ·Per-client reasoning settings tuned once and never revisited.

Effort by use case, with the outcome signal closed.

  • ·A ceiling per `x-ace-use-case` rather than per key.
  • ·Task-outcome feedback that raises the ceiling where lowered reasoning cost a retry.
team@acefleet.dev →