all docs

Model router

Intent and complexity-based model selection to optimize quality versus token cost.

problem it solves
Stops sending routine requests to your most expensive model when a cheaper model you have approved would do.

What it does

What it does: Keeps the requested model unless a model on your allowlist, on the same provider credential, is cheaper for this request and clears the quality floor for its task and difficulty.

Your allowlist: The models the router may route to. On hosted ACE you set it on this page; on-prem it lives in your ace/ config (`llm_router.allowlist`) and changes only through a reviewed PR. Each model inherits capability and price from the ACE Model Market Catalog; on-prem, a price you state overrides the catalog's.

Learning from traffic: ACE records which models your requests ask for and suggests the ones your allowlist is missing: pre-filled on this page on hosted ACE, proposed as an ace/ change on-prem.

When it keeps: Signed reasoning or a warm prompt cache on the request, a difficulty the request does not show (long multi-turn agent turns), or no cheaper allowlisted model that clears the floor.

Never refuses: A request for a model that is not on the allowlist is still served, either as sent or routed to a cheaper allowlisted model.

What we need from you

  • Your provider credential on the requestrequired

    The router chooses among models sold on the credential the request carries (x-ace-<provider>-key or your stored key), so it never switches vendor.

  • Models on your allowlistrecommended

    Choose the models the router may route to on this page (hosted), or list them under llm_router.allowlist in ace/ (on-prem). Suggestions from your traffic are pre-filled. Without a list the router picks among the catalog's models on the request's provider.

  • Request history for proposalsrecommended

    Proposals come from recorded requests, so the gateway's telemetry store must be on.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Requests go to the exact model specified by the caller.No llm_router stage recorded. Requested models are still recorded for proposals.
shadowRuns the decision and records the model it would have served without changing the request.x-ace-route-shadow-model and the shadow saving; x-ace-route-reason says why.
prodServes the chosen allowlisted model on the same provider credential.x-ace-route-model, x-ace-route-reason, x-ace-route-complexity, on streamed and buffered responses; saving logged.

Current policy

Candidate poolYour allowlist, same provider credentialNever another vendor. Without a list, the catalog's models on that provider.
Unknown difficultyKeep the requested modelMulti-turn agent turns with no structural signal are not guessed as medium.
Quality floorslow 0.55 · medium 0.65 · high 0.80High also needs a direct benchmark score for the task category.

Worth knowing before you enable it

  • ·Every allowlisted model must be in the ACE Model Market Catalog; on-prem, a model the catalog does not sell on your provider needs a stated price.
  • ·On-prem, your price wins over the catalog's; the allowlist read-back lists every difference.
  • ·On-prem, allowlist changes take effect only after the ace/ PR is reviewed and deployed.
  • ·The router respects explicit model pins (x-ace-route-to).

What it replaces

  • ·Hardcoded if/else model routing logic in microservices.
  • ·Manual prompt complexity estimation scripts.
  • ·Ad-hoc model switching SDK wrappers.

Custom routing matrices and SLA rules are available on the enterprise tier.

  • ·Custom intent taxonomy and fine-tuned classifier weights.
  • ·Latency vs cost optimization trade-off sliders.
  • ·Strict model fallback pinning per tenant.
team@acefleet.dev →