Model router
Intent and complexity-based model selection to optimize quality versus token cost.
problem it solves
Stops sending routine requests to your most expensive model when a cheaper model you have approved would do.
What it does
What it does: Keeps the requested model unless a model on your allowlist, on the same provider credential, is cheaper for this request and clears the quality floor for its task and difficulty.
Your allowlist: The models the router may route to. On hosted ACE you set it on this page; on-prem it lives in your ace/ config (`llm_router.allowlist`) and changes only through a reviewed PR. Each model inherits capability and price from the ACE Model Market Catalog; on-prem, a price you state overrides the catalog's.
Learning from traffic: ACE records which models your requests ask for and suggests the ones your allowlist is missing: pre-filled on this page on hosted ACE, proposed as an ace/ change on-prem.
When it keeps: Signed reasoning or a warm prompt cache on the request, a difficulty the request does not show (long multi-turn agent turns), or no cheaper allowlisted model that clears the floor.
Never refuses: A request for a model that is not on the allowlist is still served, either as sent or routed to a cheaper allowlisted model.
What we need from you
- Your provider credential on the requestrequired
The router chooses among models sold on the credential the request carries (x-ace-<provider>-key or your stored key), so it never switches vendor.
- Models on your allowlistrecommended
Choose the models the router may route to on this page (hosted), or list them under llm_router.allowlist in ace/ (on-prem). Suggestions from your traffic are pre-filled. Without a list the router picks among the catalog's models on the request's provider.
- Request history for proposalsrecommended
Proposals come from recorded requests, so the gateway's telemetry store must be on.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Requests go to the exact model specified by the caller. | No llm_router stage recorded. Requested models are still recorded for proposals. |
| shadow | Runs the decision and records the model it would have served without changing the request. | x-ace-route-shadow-model and the shadow saving; x-ace-route-reason says why. |
| prod | Serves the chosen allowlisted model on the same provider credential. | x-ace-route-model, x-ace-route-reason, x-ace-route-complexity, on streamed and buffered responses; saving logged. |
Current policy
| Candidate pool | Your allowlist, same provider credential | Never another vendor. Without a list, the catalog's models on that provider. |
| Unknown difficulty | Keep the requested model | Multi-turn agent turns with no structural signal are not guessed as medium. |
| Quality floors | low 0.55 · medium 0.65 · high 0.80 | High also needs a direct benchmark score for the task category. |
Worth knowing before you enable it
- ·Every allowlisted model must be in the ACE Model Market Catalog; on-prem, a model the catalog does not sell on your provider needs a stated price.
- ·On-prem, your price wins over the catalog's; the allowlist read-back lists every difference.
- ·On-prem, allowlist changes take effect only after the ace/ PR is reviewed and deployed.
- ·The router respects explicit model pins (x-ace-route-to).
What it replaces
- ·Hardcoded if/else model routing logic in microservices.
- ·Manual prompt complexity estimation scripts.
- ·Ad-hoc model switching SDK wrappers.
Custom routing matrices and SLA rules are available on the enterprise tier.
- ·Custom intent taxonomy and fine-tuned classifier weights.
- ·Latency vs cost optimization trade-off sliders.
- ·Strict model fallback pinning per tenant.