all docs

Agent trajectory router

Per-function model, effort, output-cap and cache decisions across a multi-turn agent session — never a model switch mid-session.

problem it solves
On an agent loop the prefix is re-sent every turn and prompt caches are model-scoped, so a per-request router that swaps models mid-session re-bills the whole cache. Cache misses and fixed thinking budgets then dominate the bill, not which model answered.

What it does

What it does: Reads four headers your app already sends or can send — `x-ace-use-case` (mapped to a function: `act`, `answer`, `gate`, `transform`, `judge`, `label`), `x-ace-archetype`, `x-ace-session` / `x-ace-cohort` / `x-ace-agent-depth` — and resolves a per-function policy: which tier serves it (`smart` / `fast` / `lite`), what thinking / effort / output caps apply, whether a cheap-tier answer must pass a validator, and whether the semantic cache is consulted.

Continuity guard: A model or effort change is applied only at a switch point for the scope — turn 1, prefix-TTL expiry, a trajectory-compaction epoch, a cohort's prefix version change. Anywhere else it is recorded as `deferred` with the reason and the pinned model serves. This is the guarantee the prompt model router cannot make.

Validated downgrade: `gate`, `transform`, `label` and `answer` run on the cheap tier and a validator judges the answer — JSON schema / enum / regex (mechanical), citation containment (grounded), or a `judge` call on a different model family (rubric). A failure re-dispatches on the function's `escalate_to` tier and the response says `x-ace-route-escalated: true`. A per-function escalation-rate breaker reverts the function to its flagship tier before a bad downgrade can double cost.

Cache-stability audit (shadow-only): Fingerprints the prefix between consecutive requests in a scope and names the first divergent block and its cause — `timestamp`, `conditional_section`, `tool_list_churn`, `retrieval_reorder`, `single_tail_breakpoint` — on the trace. It never rewrites anything; it tells you why your cache reads are low.

Relationship to the prompt model router: `llm_router` stays unchanged and becomes this skill's fallback for `continuity: none` traffic (pipelines, one-shot generation, untagged requests), where per-request intent classification is the right tool. Inside a pinned session it is not consulted.

What we need from you

  • At least one provider key in Provider Vaultrequired

    Tiers resolve to concrete models through the Model Market catalog and need stored keys to dispatch to them.

  • `x-ace-use-case` on every requestrequired

    The purpose tag is what a function policy keys on. Untagged requests are never purpose-routed; they fall through to `llm_router` or the caller's model.

  • `x-ace-session` (or `x-ace-cohort`) on agent trafficrecommended

    The scope the continuity guard pins to. Without one the request is `continuity: none` and the guard has nothing to hold.

  • A tenant policy naming tiers and use-case → function mappingsrecommended

    Archetype presets (`agentic`, `qa`, `generative`, `pipeline`, `orchestrator`) and function defaults are built in; the only vertical-specific part is the `use_cases` map.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Requests go to the caller's model or to whatever `llm_router` decides.No agent_trajectory_router stage recorded.
shadowResolves the policy, runs the continuity guard, the shape rewrites and the validators, then restores the body and sends it as received. The cache-stability audit runs here.Stage with function, tier, would_route, would_save_usd, deferred_reason, shape_rewrites and cache_stability; the per-function spend ledger fills in even when nothing else is on.
prodRewrites model, effort and output controls per the resolved policy at switch points; serves the pinned model elsewhere; escalates a failed cheap-tier answer.x-ace-route-model / -tier / -reason on every routed response; x-ace-route-escalated on a re-dispatch; x-ace-warnings: agent_trajectory_router_reverted when the breaker trips.

Current policy

Tierssmart: sonnet-current · fast: haiku-current, flash-current · lite: flash-lite-currentAliases over your enabled models; a tier can be a ranked list and the breaker picks within it.
Function tiersact → smart (pinned) · answer → fast+grounded · transform → fast+mechanical · gate, label → lite · judge → fast, never the producerEach has an `escalate_to`; `act` is never downgraded, only shaped.
Switch pointssession: turn 1, prefix expiry, compaction epoch · cohort: prefix version change, sticky canary bucket · lineage: inherit parent's pinOutside these the decision is `deferred` and the pinned model serves.
Escalation breaker20% escalation rate over 60 s (min 5 samples), or 5 consecutive failures; 30 s cooldown, backing off to 300 sTrips per tenant × function; the function reverts to its `escalate_to` tier while open.
Semantic cacheon for answer / gate in QA and pipeline archetypes; off for act / generateUnique agent-turn prompts never hit; the lookup is skipped rather than wasted.

Worth knowing before you enable it

  • ·A caller's `x-ace-route-to` always wins, and so does a replayed signed thinking block — the router stands down rather than break a provider's verification.
  • ·This skill does not switch reasoning on and does not raise an output cap; it lowers, pins and holds. If you want a hard per-key ceiling regardless of function, that is `output_budget` / `reasoning_effort`, which run after it.
  • ·The scope store (shared with agent trajectory compaction) being unavailable is treated as `continuity: none` but the router still never switches — the safe default is no change.
  • ·`acceptance` validators (`x-ace-task-status`, `x-ace-user-feedback`) arrive one request late, so they feed the escalation statistics, not same-request escalation.
  • ·Fail-open: an exception in policy resolution relays the request untouched and reports `pipeline_bypassed`.
  • ·The trace stage is also emitted under the legacy id `model_router`; both spellings are accepted in `x-ace-skills` and the policy body.

What it replaces

  • ·`getModel('smart' | 'fast')` calls scattered across every agent call site.
  • ·Per-tenant model-override plumbing threaded through orchestrators and tools.
  • ·The 'add a purpose column to the usage table' migration — `x-ace-use-case` already lands on the request row.
  • ·Hand-rolled 'don't switch models mid-session' guards and cache-warm bookkeeping.

Per-tenant routing policies, custom validators and the cache-stability audit in production are on the enterprise tier.

  • ·Custom tier definitions and per-use-case overrides beyond the built-in archetypes.
  • ·Grounded and rubric validators against your own sources and judge models.
  • ·Per-function spend ledger and cache-stability panel on the console.
team@acefleet.dev →