Model routing is the highest-leverage cost decision in an LLM stack. Most production traffic does not need a frontier model, and a router that sends simple queries to a cheap endpoint while escalating the hard ones keeps quality where it matters and cuts the bill everywhere else. The evidence is concrete: in the RouteLLM benchmark work, routers trained on preference data kept 95% of GPT-4 quality on MT Bench while cutting cost by up to 3.66x, and the paper reports those routers transfer to model pairs they were never trained on. This guide turns that result into an implementable design: a decision surface you own, a pricing model you recompute monthly, and guardrails that catch quality regressions before your customers do.
What Routing Actually Decides
At its core, query routing is a binary classification problem: given an incoming prompt, decide whether a strong model or a weak model should answer it. The RouteLLM framework formalizes this as a routing function with a single cost threshold: the router predicts the probability that the strong model would win a head-to-head comparison, and routes to the strong model only when that probability clears the threshold. Raising the threshold sends more traffic to the cheap model; lowering it buys quality at premium rates.
Three engineering realities follow. First, the router itself must be cheap and fast — an embedding-based classifier or a small BERT-style model adds milliseconds, while calling another LLM to classify adds both latency and its own token spend. Second, routing decisions are per-query, not per-user or per-session, so your telemetry needs query-level granularity. Third, the threshold is a business dial, not a technical constant: product surfaces with low error tolerance (legal drafting, code that ships) should run a stricter threshold than internal tooling.
The Pricing Gap You Exploit
Routing works because the price spread between model tiers is enormous. In the RouteLLM cost model, GPT-4 averaged $24.7 per million tokens while Mixtral 8x7B averaged $0.24, a gap of roughly two orders of magnitude. When the strong model costs on the order of a hundred times more than the weak one, even a conservative router that escalates half of traffic still halves spend — and every percentage point of traffic you keep on the cheap tier compounds monthly.
Modern API pricing adds a second lever: time-of-day rates. DeepSeek’s published API pricing is the clearest example — off-peak rates are exactly half of peak, and peak windows run 01:00 – 04:00 and 06:00 – 10:00 UTC on weekdays. A European team can batch-classify, summarize, and backfill embeddings outside those windows and pay half price for identical tokens, before any routing decision is even made. Combine the two: schedule batchable work off-peak, and reserve peak hours plus premium models for interactive traffic that actually needs them.
| Query signal | Default route | Escalate when | Why it saves |
|---|---|---|---|
| Short factual or formatting request | Small model | Retry fails or user reformulates | Highest volume tier, rarely needs frontier capability |
| Multi-step reasoning or analysis | Frontier model | Not applicable | Wrong-answer cost exceeds token cost |
| Long-document summarization | Mid-tier model with prefix caching | Summary misses stated constraints | Cache hits cut repeated input cost dramatically |
| Code generation with tests | Mid-tier model, then repair loop | Tests fail twice | Repair loops are cheaper than always paying premium |
| Embedded classification or extraction | Small model, batch off-peak | Schema violations spike | Structured outputs tolerate weaker models |
Building the Router: Four Steps
Start simple and earn complexity. A cold-start router built from rules captures most of the savings with none of the training cost.
- Log and label your traffic. Sample real prompts and run them through both the strong and weak model. Label which answers win, tie, or lose. This preference log is the ground truth every later decision depends on, and it is the only dataset that reflects your actual distribution instead of public benchmarks.
- Route by rules first. Prompt length, task type keywords, language, and presence of structured output requests classify a surprising share of traffic. Route the obviously cheap tier now; you do not need ML to know that a 40-token formatting request does not need a frontier model.
- Add a learned router only where rules fail. Train an embedding classifier on your preference log, calibrate the escalation threshold against a held-out set, and ship it behind the same interface so you can A/B rule-based against learned routing.
- Wire fallbacks and cooldowns. Routing layers like LiteLLM handle the operational side: its least-busy strategy sends each request to the deployment with the fewest active requests, and ordered deployment fallbacks move traffic to the next tier or provider on connection errors and rate limits, with cooldowns keeping a failing endpoint out of rotation instead of poisoning every retry.
Guardrails That Keep Quality
A router is a bet that the cheap model is good enough, and bets need monitoring. Keep a frozen canary evaluation set of a few hundred real prompts with known-good answers, and run it through every routing configuration change before rollout — the same discipline that applies to LLM evaluation pipelines in CI. Alert on escalation rate, not just spend: a sudden jump in escalations means your cheap tier is silently degrading and the router is papering over it at premium prices.
Watch drift in both directions. Query distributions shift as your product changes, and model providers update underlying weights without notice — a threshold calibrated in one quarter can misroute the next. Cluster your traffic periodically, the same way you would when choosing embedding models for semantic search, so new query categories get an explicit routing rule instead of falling through to a default. Finally, log the model, threshold score, and cost of every served request; you cannot optimize a routing decision you cannot reconstruct.