Agent trajectory router
Per-function model, effort, output-cap and cache decisions across a multi-turn agent session — never a model switch mid-session.
problem it solves
On an agent loop the prefix is re-sent every turn and prompt caches are model-scoped, so a per-request router that swaps models mid-session re-bills the whole cache. Cache misses and fixed thinking budgets then dominate the bill, not which model answered.
What it does
What it does: Reads four headers your app already sends or can send — `x-ace-use-case` (mapped to a function: `act`, `answer`, `gate`, `transform`, `judge`, `label`), `x-ace-archetype`, `x-ace-session` / `x-ace-cohort` / `x-ace-agent-depth` — and resolves a per-function policy: which tier serves it (`smart` / `fast` / `lite`), what thinking / effort / output caps apply, whether a cheap-tier answer must pass a validator, and whether the semantic cache is consulted.
Continuity guard: A model or effort change is applied only at a switch point for the scope — turn 1, prefix-TTL expiry, a trajectory-compaction epoch, a cohort's prefix version change. Anywhere else it is recorded as `deferred` with the reason and the pinned model serves. This is the guarantee the prompt model router cannot make.
Validated downgrade: `gate`, `transform`, `label` and `answer` run on the cheap tier and a validator judges the answer — JSON schema / enum / regex (mechanical), citation containment (grounded), or a `judge` call on a different model family (rubric). A failure re-dispatches on the function's `escalate_to` tier and the response says `x-ace-route-escalated: true`. A per-function escalation-rate breaker reverts the function to its flagship tier before a bad downgrade can double cost.
Cache-stability audit (shadow-only): Fingerprints the prefix between consecutive requests in a scope and names the first divergent block and its cause — `timestamp`, `conditional_section`, `tool_list_churn`, `retrieval_reorder`, `single_tail_breakpoint` — on the trace. It never rewrites anything; it tells you why your cache reads are low.
Relationship to the prompt model router: `llm_router` stays unchanged and becomes this skill's fallback for `continuity: none` traffic (pipelines, one-shot generation, untagged requests), where per-request intent classification is the right tool. Inside a pinned session it is not consulted.
What we need from you
- At least one provider key in Provider Vaultrequired
Tiers resolve to concrete models through the Model Market catalog and need stored keys to dispatch to them.
- `x-ace-use-case` on every requestrequired
The purpose tag is what a function policy keys on. Untagged requests are never purpose-routed; they fall through to `llm_router` or the caller's model.
- `x-ace-session` (or `x-ace-cohort`) on agent trafficrecommended
The scope the continuity guard pins to. Without one the request is `continuity: none` and the guard has nothing to hold.
- A tenant policy naming tiers and use-case → function mappingsrecommended
Archetype presets (`agentic`, `qa`, `generative`, `pipeline`, `orchestrator`) and function defaults are built in; the only vertical-specific part is the `use_cases` map.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Requests go to the caller's model or to whatever `llm_router` decides. | No agent_trajectory_router stage recorded. |
| shadow | Resolves the policy, runs the continuity guard, the shape rewrites and the validators, then restores the body and sends it as received. The cache-stability audit runs here. | Stage with function, tier, would_route, would_save_usd, deferred_reason, shape_rewrites and cache_stability; the per-function spend ledger fills in even when nothing else is on. |
| prod | Rewrites model, effort and output controls per the resolved policy at switch points; serves the pinned model elsewhere; escalates a failed cheap-tier answer. | x-ace-route-model / -tier / -reason on every routed response; x-ace-route-escalated on a re-dispatch; x-ace-warnings: agent_trajectory_router_reverted when the breaker trips. |
Current policy
| Tiers | smart: sonnet-current · fast: haiku-current, flash-current · lite: flash-lite-current | Aliases over your enabled models; a tier can be a ranked list and the breaker picks within it. |
| Function tiers | act → smart (pinned) · answer → fast+grounded · transform → fast+mechanical · gate, label → lite · judge → fast, never the producer | Each has an `escalate_to`; `act` is never downgraded, only shaped. |
| Switch points | session: turn 1, prefix expiry, compaction epoch · cohort: prefix version change, sticky canary bucket · lineage: inherit parent's pin | Outside these the decision is `deferred` and the pinned model serves. |
| Escalation breaker | 20% escalation rate over 60 s (min 5 samples), or 5 consecutive failures; 30 s cooldown, backing off to 300 s | Trips per tenant × function; the function reverts to its `escalate_to` tier while open. |
| Semantic cache | on for answer / gate in QA and pipeline archetypes; off for act / generate | Unique agent-turn prompts never hit; the lookup is skipped rather than wasted. |
Worth knowing before you enable it
- ·A caller's `x-ace-route-to` always wins, and so does a replayed signed thinking block — the router stands down rather than break a provider's verification.
- ·This skill does not switch reasoning on and does not raise an output cap; it lowers, pins and holds. If you want a hard per-key ceiling regardless of function, that is `output_budget` / `reasoning_effort`, which run after it.
- ·The scope store (shared with agent trajectory compaction) being unavailable is treated as `continuity: none` but the router still never switches — the safe default is no change.
- ·`acceptance` validators (`x-ace-task-status`, `x-ace-user-feedback`) arrive one request late, so they feed the escalation statistics, not same-request escalation.
- ·Fail-open: an exception in policy resolution relays the request untouched and reports `pipeline_bypassed`.
- ·The trace stage is also emitted under the legacy id `model_router`; both spellings are accepted in `x-ace-skills` and the policy body.
What it replaces
- ·`getModel('smart' | 'fast')` calls scattered across every agent call site.
- ·Per-tenant model-override plumbing threaded through orchestrators and tools.
- ·The 'add a purpose column to the usage table' migration — `x-ace-use-case` already lands on the request row.
- ·Hand-rolled 'don't switch models mid-session' guards and cache-warm bookkeeping.
Per-tenant routing policies, custom validators and the cache-stability audit in production are on the enterprise tier.
- ·Custom tier definitions and per-use-case overrides beyond the built-in archetypes.
- ·Grounded and rubric validators against your own sources and judge models.
- ·Per-function spend ledger and cache-stability panel on the console.