Survey · 2026-10-03 · Model Intelligence / Faros
Why Learned AI Model Routing Must Beat Good Rules
A learned router that beats "always use the most expensive model" has shown that cheaper models exist. It has not shown that learning was worth it.
Abstract
Learned LLM routing promises to send each request to the cheapest model that will handle it well. A large literature now proposes routers trained on preference data, difficulty predictions, cascades and bandit feedback. This survey asks a narrower question: what evidence would show that a learned router adds value over a well-designed rule-based policy, and does the published evidence meet that standard? We organize routing methods into four families and summarize what standardized benchmarks have found. We note in particular the result that several recent approaches, including commercial routers, fail to reliably outperform a simple baseline under unified evaluation, and the identification of routing collapse in score-predicting routers. We then argue that strong static rules, not a frontier-only policy, are the correct baseline. We describe how caching can masquerade as routing intelligence, how logged data can bias router training, and how non-inferiority testing should govern promotion. The survey reports no Ethen measurements. It motivates the protocol in a companion paper.
The question
Every routing paper compares its method with something. The choice of that something largely determines how impressive the result looks. Compared with sending every request to the most capable and most expensive model, almost any router that sometimes uses a cheaper model will save money at small quality cost, because many requests are easy. That comparison establishes that tiering is useful. It says little about whether a learned decision rule is better than a simple designed one, such as "use the small model for classification and extraction, the large model for multi-step reasoning, and the reasoning model for anything with a failing test".
For a system that will actually deploy a router, the second question is the one that matters. Learned routers carry costs that rules do not: training data, retraining as models and prices change, monitoring for drift, debugging opaque decisions, and the risk of exploration. A learned router is worth those costs only if it beats the best simple alternative.
Four families of routing methods
Figure 1 organizes the literature.
Figure 1. A taxonomy of LLM routing approaches. Four families of routing methods in the literature, distinguished by when the decision is made and what signal it uses. Examples are representative, not exhaustive. Evidence label: TAXONOMY. Source: Survey of cited literature.
Static rules. Deterministic policies keyed on task type, required capabilities, cost tier, latency budget and constraints such as data residency. Rules are cheap, auditable and easy to change. They are rarely the subject of papers, which is part of the problem: they are often omitted as baselines.
Predictive routers decide which model to use before any generation, based on features of the request. Hybrid LLM trains a router to choose between a small and a large model according to predicted query difficulty and a tunable quality threshold (Ding et al.). RouteLLM learns routers from human preference data, framing the decision as whether a strong model is needed (Ong et al.). RouterBench standardizes evaluation of such routers across models and tasks (Hu et al.).
Cascades decide after a cheap attempt. FrugalGPT showed that a learned cascade over several models can match the best individual model at substantially lower cost on the tasks studied (Chen et al.). AutoMix uses the small model's self-verification to decide whether to escalate (Aggarwal et al.).
Online and bandit routers learn from feedback as they operate. Budget-constrained contextual bandits are a natural formulation when costs matter and labels arrive over time (Panda et al.).
What standardized evaluation has found
Two recent results reframe the field.
LLMRouterBench assembled more than 400,000 instances from 21 datasets and 33 models and re-evaluated routing methods under one framework. It confirmed strong complementarity among models, which is the premise of routing. It also found that many routing methods perform similarly, that several recent approaches including commercial routers fail to reliably outperform a simple baseline, and that a substantial gap to the oracle remains, driven mainly by failures to recall which model would have succeeded [EXTERNAL PRIMARY-SOURCE RESULT] (Li et al.). Two further findings matter for practice: the choice of embedding backbone had limited impact, and larger model ensembles gave diminishing returns compared with careful curation of a few models.
Routing collapse. A separate study found that as the cost budget increases, routers that predict scalar performance scores systematically default to the most capable and most expensive model, even when cheaper models suffice. The authors attribute this to a mismatch between the training objective, predicting scores, and the decision, a discrete comparison among candidates. Small prediction errors can flip orderings. A router that learns rankings directly mitigated the effect on RouterBench (Lai et al.).
A third result is theoretical. In a game between a provider with two models and a user who can re-prompt or abandon, the provider's optimal policy was almost always static, with no cascading. There was a misalignment between provider-optimal and user-preferred routes when the two rank models differently (Mahmood). The model is stylized, but it cautions against assuming that dynamic escalation always pays.
Taken together, these findings support a sober reading. Routing is valuable, because models are complementary. Learning to route well is hard, gains over simple methods are often small, and the headroom to the oracle is real but not easily captured.
The right baseline
Figure 2 summarizes what each baseline comparison can show.
Figure 2. What each baseline comparison can show. The baseline determines what a routing result means. Only the comparison against strong rules isolates the value of learning; frontier-only comparisons mostly measure the value of using cheaper models at all. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis of cited literature.
A frontier-only baseline mainly measures the value of using cheaper models at all. A cheapest-only baseline mainly measures the value of ever escalating. An oracle, the best model per query in hindsight, measures headroom and is useful as an upper bound, but it is not deployable. Only a strong static rule set, written by people who understand the workload and given the same information the learned router sees, isolates the incremental value of learning.
Ethen's internal reconciliation work adopted this standard explicitly. Its proposed promotion gate for any learned routing stage is measured against good rules, not against frontier-only. Savings versus frontier-only may be reported as secondary context, but never as the justification [PROPOSED TARGET]. We recommend that routing papers do the same: include a strong rule baseline, describe how it was constructed, and report results against it first.
What makes rules "strong"? At minimum: they are keyed on the same task features available to the learned router; they are tuned on the same development data; they encode hard constraints such as capability limits, context length and residency; and they include a sensible escalation path for failures. A rule set written in an afternoon without access to development data is a weak baseline, and beating it is weak evidence.
Confounds that inflate routing gains
Caching. Prompt caching makes repeated prefixes cheaper. A router that tends to keep a conversation on one provider benefits from warm caches; one that switches providers pays to rebuild them. Cost differences between routers can therefore reflect cache placement rather than decision quality. A factorial design, crossing rules or learned policy with cache-aware or cache-blind placement, separates the two.
Pricing changes. Provider prices change frequently. A router evaluated before and after a price change can appear to improve or degrade without any change in its decisions. Comparisons should use prices frozen at a stated date, with sensitivity analysis.
Selection bias in training data. A router trained on logs from a previous policy sees outcomes only for the models that policy chose. If the previous policy always sent hard requests to the large model, the logs contain almost no evidence about how the small model handles hard requests. Unbiased estimation of a new policy's value from such logs requires the selection probabilities of the logging policy (Li et al., 2010; Dudík et al.). See Why AI Routers Should Log Propensities From Day One.
Benchmark-workload mismatch. Routing benchmarks draw on academic datasets. Enterprise agent workloads involve multi-step tasks with tools, where the unit of success is a verified outcome rather than a correct answer. Gains measured per query may not survive aggregation to tasks, especially if a cheap model's failures trigger retries.
The objective: cost per verified outcome
Routing studies commonly report cost and accuracy separately or along a Pareto frontier. For agents that perform work, we propose a single primary objective: cost per verified outcome, all eligible delivery costs divided by the number of independently verified successful outcomes, with quality protected by a separate constraint. The objective naturally penalizes routers that save on the first call but trigger retries, escalations or failures later. It is developed in Cost Per Verified Outcome. A related proposal in the evaluation literature, measuring the expected cost of obtaining a correct solution, makes a similar argument for language models generally (Erol et al.). The broader case for jointly reporting cost and accuracy in agent evaluation is made by Kapoor et al.
Non-inferiority, not "no significant difference"
A router that saves money is useful only if it does not lose too much quality. "No significant difference in quality" is not evidence of no loss: with small samples, large losses go undetected. The correct framing is non-inferiority. Before the experiment, choose a margin, such as two percentage points, and show that the lower confidence bound on the quality difference lies above minus that margin. Two one-sided tests provide the standard procedure (Schuirmann). Online controlled-experiment practice adds further guidance on sample-ratio checks, variance reduction and guarding against peeking (Kohavi et al.).
Two statistical facts shape what is achievable. First, narrow margins require large samples. At success rates near 80%, the standard error of a difference between two independent proportions with 500 tasks each is roughly 2.5 percentage points, so a margin of half a point cannot be demonstrated at that scale [ILLUSTRATIVE EXAMPLE — arithmetic, not measured]. Second, tasks cluster by family and source, so naive per-task confidence intervals are too narrow. Analyses should cluster or pair by task family.
Figure 3 shows the two-condition promotion rule we recommend.
Figure 3. A promotion rule with two conditions. Proposed decision rule for promoting a learned router over rules. Cost must fall by a pre-registered margin and quality must be shown non-inferior, meaning the lower confidence bound on the quality difference clears the margin. Thresholds shown are Ethen proposed targets. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen proposed targets; not measured.
When rules win
Rules winning is a legitimate and useful result. It means the workload's structure is simple enough to capture by hand, or that the learned router lacks the features or data to do better. In either case, the right response is to keep rules, publish the comparison, and revisit when the data or the workload changes. A learned router is a depreciating asset in a non-stationary environment: models, prices and provider quality change monthly. Even a router that wins today needs ongoing evaluation to keep winning.
There is also a middle path. A learned component can maintain a rule set: proposing threshold changes when model prices or capabilities shift, for humans to review. That captures some adaptivity without handing production decisions to an opaque policy.
Position
We draw five recommendations from the literature:
- Report results against a strong rule baseline first, and describe how it was built.
- Separate cache and pricing effects from decision effects with factorial designs.
- Log selection probabilities from the first routing decision so that off-policy evaluation is possible later.
- Use cost per verified outcome as the primary objective for agent workloads.
- Promote learned routers only on non-inferiority evidence with a pre-registered margin.
The companion protocol, How to Test Whether Learned AI Routing Beats Strong Rules, turns these recommendations into an experimental design. The broader decision layer they serve is described in Faros: Researching How Intelligence Should Choose Intelligence. How much routing value survives falling prices is discussed in Why Better Foundation Models May Make Evaluation More Valuable, Not Less.
Limitations
This survey covers a fast-moving literature selectively and relies on benchmark results whose task distributions differ from enterprise agent workloads. We report published findings at the level of their abstracts and main claims; we have not reproduced them. Our recommended baseline standard is a judgment, and reasonable researchers may weigh the costs of learned routers differently. No Ethen routing data informs this survey.
Conclusion
Routing among models is useful because models differ. Learning to route is useful only if it beats what a careful engineer would write by hand, after accounting for caching, pricing and selection bias, on the outcomes that matter. The published evidence suggests that bar is often not met. That is not a reason to stop researching learned routing. It is a reason to hold it to the right standard.
FAQ
Are learned routers worse than rules? Not necessarily. The evidence is that they often fail to reliably beat simple baselines under unified evaluation. Whether a learned router beats strong rules depends on the workload and must be tested.
Why not compare against the most capable model? Because nearly any tiering beats it on cost. That comparison shows cheaper models are often good enough, not that learning helps.
What is non-inferiority testing? A test showing that a new method's quality is not worse than a baseline's by more than a pre-specified margin, using the lower confidence bound on the difference.
Related research
- Faros: Researching How Intelligence Should Choose Intelligence — the broader decision layer.
- How to Test Whether Learned AI Routing Beats Strong Rules — the experiment protocol this survey motivates.
- Why AI Routers Should Log Propensities From Day One — selection bias in router training data.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI — the economic objective.
- Why Better Foundation Models May Make Evaluation More Valuable, Not Less — what routing value survives cheaper models.
References
- Ding, D. et al. (2024). Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing. arXiv:2404.14618. https://arxiv.org/abs/2404.14618
- Ong, I. et al. (2024). RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665. https://arxiv.org/abs/2406.18665
- Hu, Q. J. et al. (2024). RouterBench: A Benchmark for Multi-LLM Routing System. arXiv:2403.12031. https://arxiv.org/abs/2403.12031
- Chen, L., Zaharia, M., Zou, J. (2023). FrugalGPT. arXiv:2305.05176. https://arxiv.org/abs/2305.05176
- Aggarwal, P. et al. (2023). AutoMix: Automatically Mixing Language Models. arXiv:2310.12963. https://arxiv.org/abs/2310.12963
- Panda, P. et al. (2025). Adaptive LLM Routing under Budget Constraints. arXiv:2508.21141. https://arxiv.org/abs/2508.21141
- Li, H. et al. (2026). LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. arXiv:2601.07206. https://arxiv.org/abs/2601.07206
- Lai, G. et al. (2026). When Routing Collapses: On the Degenerate Convergence of LLM Routers. arXiv:2602.03478. https://arxiv.org/abs/2602.03478
- Mahmood, R. (2026). Routing, Cascades, and User Choice for LLMs. arXiv:2602.09902. https://arxiv.org/abs/2602.09902
- Li, L. et al. (2010). Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms. arXiv:1003.5956. https://arxiv.org/abs/1003.5956
- Dudík, M., Langford, J., Li, L. (2011). Doubly Robust Policy Evaluation and Learning. arXiv:1103.4601. https://arxiv.org/abs/1103.4601
- Erol, M. H. et al. (2025). Cost-of-Pass: An Economic Framework for Evaluating Language Models. arXiv:2504.13359. https://arxiv.org/abs/2504.13359
- Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
- Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. J. Pharmacokinet. Biopharm. 15:657–680. https://doi.org/10.1007/BF01068419
- Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press. https://doi.org/10.1017/9781108653985
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Faros: Researching How Intelligence Should Choose Intelligence
A position paper reframing AI model routing as an execution-configuration decision across model, context, tools, verification, recovery, cost and risk.
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work
A research note on model change assurance: replaying an organization's own historical tasks to find regressions before an LLM upgrade reaches real work.
- How to Test Whether Learned AI Routing Beats Strong Rules
A research protocol for an LLM routing experiment: strong-rule baselines, propensity-logged data, factorial cache controls, non-inferiority tests and decision rules.
Explained on the Ethen Blog
- How Ethen Gateway Chooses an Eligible Model
Eligibility is elimination: the Gateway discards every model that cannot serve a request before it ranks anything at all.
- How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading Paths
The fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.
- Why Ethen Keeps Research Separate From Product Claims
Ethen keeps research and product claims apart because they rest on different kinds of evidence, and readers make different decisions based on them. A product claim says what an Ethen product does today, and should be backed by product evidence. A research claim says what Ethen Research Lab is studying, proposing or has measured, and carries an explicit evidence status — often "proposal" or "protocol; not yet run". When the two blur, research gets read as a shipped feature, and product statements borrow credibility from papers that never tested them. So we publish in three lanes — the Ethen Blog, Ethen Research Lab and company writing — and we follow a few simple rules whenever one lane draws on another.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.