Model routing: lower cost without hiding quality failures
Use measurable task categories, bounded escalation, and explicit fallback rules to decide which model should handle each request.

A product does not always need the same model for every request. Extracting a short label and resolving an ambiguous support case are different workloads. Routing can assign them to different paths, but the router becomes another component that can make mistakes.
Evaluate the entire routing policy. A cheaper first attempt is not necessarily a cheaper completed task if it creates retries, escalations, and additional review. The savings should remain visible after every step of the workflow is counted.
Begin with observable categories
Use categories the application can identify reliably: input modality, document length, required tool, or explicit task type. Avoid starting with a vague “easy or hard” label that has no tested definition. A request can be short and still require difficult reasoning.
Imagine a catalog assistant that handles both attribute extraction and product compatibility questions. Extraction from a supported template can have deterministic validators. Compatibility may need several sources and clarification. Those differences provide a clearer routing basis than prompt length alone.
Evaluate each branch and the routing decision
Create reference cases for each task type, then include ambiguous cases near the boundaries. Check whether the selected path can complete the job and whether the router sends too much work to an expensive fallback. Compare against a simpler baseline that uses one model or one fixed workflow.
Report quality for the complete system, not only for each model when it receives ideal inputs. The router may send the wrong cases to an otherwise capable branch. Keep an explanation of the selected route in diagnostic logs without exposing private content unnecessarily.
Bound escalation
Set a maximum number of attempts and a total budget per user operation. If a first response fails validation, decide whether the next step is a retry, a different model, a request for evidence, or a manual handoff. A stronger model cannot recover facts that no branch has permission or ability to retrieve.
For an illustrative cost calculation, suppose the first step costs $0.002 per request and 20 percent of requests add a $0.020 fallback. Expected model cost is $0.006 per request before other charges: $0.002 plus 0.2 times $0.020. These are invented arithmetic inputs, not current provider prices. Use observed usage and actual rates for your product.
Distinguish failure from poor output
A timeout and a valid but unsupported answer require different handling. For transient service errors, use a bounded retry policy with backoff. For missing evidence, repeating the same request may only produce a differently worded guess.
Before switching providers, confirm that the fallback can accept the required input and follow the same data-handling constraints. Normalize response handling carefully, while retaining provider-specific errors for diagnosis. Do not mark an incomplete stream as success just because a fallback can produce a fluent continuation.
Roll out with a visible comparison
Test the policy offline, then use a limited rollout with explicit quality and cost checks. Watch task completion, corrections, tail latency, and fallback frequency. Anthropic's agent evaluation guidance emphasizes measuring systems that act across multiple steps; the same discipline is useful when a router creates several possible execution paths.
Keep a way to restore the simpler baseline. If the router saves token cost but increases abandoned tasks or support work, revise the decision rules. Routing earns its complexity when the finished workflow delivers a measured improvement that survives realistic traffic.