Research Protocol · 2026-10-03 · Model Intelligence / Faros
How to Test Whether Learned AI Routing Beats Strong Rules
A protocol for the experiment we intend to run before trusting a learned router: what is compared, how, at what thresholds, and what result would make us keep the rules.
Abstract
Published evidence shows that learned routers among language models often fail to reliably outperform simple baselines under unified evaluation, that routers trained to predict scalar scores can collapse onto the most expensive model, and that apparent routing gains can come from caching and pricing rather than from better decisions. This protocol specifies an LLM routing experiment designed to answer one question for agent workloads: does a learned execution-configuration policy reduce cost per verified outcome relative to strong rules, without a material loss of quality? The comparator is a rule set built by an independent team on development data. The design proceeds through four gated phases: offline estimation from propensity-logged data and paired replay, shadow decisions, and a randomized online comparison with bounded exposure. A factorial component separates decision quality from cache placement, and a durability factor tests whether gains survive a model change. Primary and secondary metrics, the non-inferiority analysis, power planning, stopping rules, threats to validity and decision rules are fixed in advance. The protocol has not been run. A working title that implied results was replaced with this protocol title because no measured results exist.
Background and rationale
Routing among models of different cost and capability is an established cost lever, studied through predictive routers (Ding et al.; Ong et al.), cascades (Chen et al.; Aggarwal et al.) and bandits (Panda et al.). Three findings motivate a careful protocol. A large-scale re-evaluation found that many routing methods perform similarly and that several, including commercial routers, fail to reliably beat a simple baseline (Li et al.). Routers trained to predict scalar performance scores can default to the most expensive model as budgets rise, a failure mode called routing collapse (Lai et al.). And a game-theoretic analysis found that a provider's optimal routing policy is almost always static in the setting studied (Mahmood). The case for holding learned routing to a strong-rules standard is developed in Why Learned AI Model Routing Must Beat Good Rules. The broader decision layer, which chooses context, tools, verification and recovery as well as models, is described in Faros. This protocol tests the narrowest version of the question first.
Objectives and hypotheses
Primary objective. Determine whether a learned policy achieves lower cost per verified outcome (CPVO) than strong rules on agent tasks, with quality non-inferior.
H1 (primary). The learned policy's CPVO is at least 15% lower than strong rules, and the one-sided 95% lower confidence bound on the difference in verified success rate (learned minus rules) is above −2 percentage points. These thresholds are [PROPOSED TARGET]s from Ethen's internal reconciliation work. They are fixed before the study and revisited only through a recorded amendment before any outcome data are seen.
H2 (decomposition). The CPVO reduction persists in the cache-blind arms, so it is not explained by cache placement alone.
H3 (durability). After a scheduled model change during the study, the learned policy's advantage persists without re-fitting, or is restored within a stated re-fitting budget.
Estimands
- Primary: the ratio of CPVO under the learned policy to CPVO under strong rules, over the task population defined below, at the task-success verification level.
- Co-primary quality estimand: the difference in verified success rates, learned minus rules.
- Secondary: differences in latency (median and 95th percentile), escalation rate, critical-incident rate, reliability across repeated trials (pass^k; Yao et al.) and decision overhead (the cost of computing the decision itself).
CPVO follows the definition in Cost Per Verified Outcome. It includes failed attempts and verification costs, priced at frozen, dated rates.
Population and tasks
Tasks are drawn from at least two families with strong verification, such as code maintenance with test suites and support operations with state-based checks. Families are fixed in advance and the analysis is stratified by family. Tasks must carry rights permitting their use in this research. A pilot sample, used only for variance estimation and power planning, is excluded from the confirmatory analysis.
Arms
- Strong rules (primary comparator). A deterministic policy keyed on task features, capabilities, context length, cost tier, residency and provider health, with stickiness and a fallback path. It is built by a team that does not build the learned policy, tuned on the same development data, and frozen before Phase 1.
- Learned policy. A ranking model over feasible configurations, trained on development data. It ranks rather than predicting scalar scores, given the routing-collapse evidence.
- Nearest-neighbor router. A simple learned baseline that chooses the configuration that worked best on the most similar past tasks.
- Frontier-only (secondary reference). Always the most capable configuration. It is reported for context and never used as the gate.
- Oracle (diagnostic). The best configuration per task in hindsight, computed from replay. It shows headroom and is not deployable.
All arms share the same feasible set: no arm may use a configuration that the governing mandate or tenant policy forbids.
Phased design
Figure 1 shows the four phases.
Figure 1. Four phases, each gating the next. The protocol moves from cheap, risk-free evidence to costly, live evidence only when the cheaper phase is favorable. Rules remain the production policy until the online phase passes its pre-registered gate. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Phase 0: build the comparator. The rules team writes and tunes strong rules on development data, documents them, and freezes them.
Phase 1: offline evaluation. Two sources of counterfactual evidence are used. First, logged decisions from a period in which a small, mandate-bounded share of eligible decisions was randomized among feasible configurations, with selection probabilities recorded. Inverse propensity scoring and doubly robust estimators give unbiased estimates of each policy's value under stated assumptions (Li et al., 2010; Dudík et al.), with weight clipping and effective-sample-size reporting. The logging requirement is explained in Why AI Routers Should Log Propensities From Day One. Second, paired replay of a task sample under each arm's chosen configuration, in environments validated for replay fidelity, as described in Counterfactual Replay for AI Agents. Phase 1 passes if the learned policy's estimated CPVO advantage has an interval excluding zero and its quality estimate is not clearly inferior.
Phase 2: shadow. The learned policy makes decisions on live traffic that are recorded but not executed. Its choices are compared with the rules' choices and with Phase 1 predictions. Phase 2 passes if shadow decisions are consistent with the offline estimates and no hard-constraint violations occur.
Phase 3: online randomized comparison. Eligible tasks are randomized between rules and learned policy, with exposure bounded by risk class. Only low-risk task classes enter initially, and high-risk classes are excluded unless explicitly approved. Randomization is at the task level with stickiness within a task, following standard online-experimentation practice for unit-of-randomization choice (Kohavi et al.), because switching models mid-task confounds attribution and destroys cache benefits.
Factorial component
Figure 2 shows the factorial design used within Phase 1 replay and, where feasible, Phase 3.
Figure 2. Factorial design separating decision quality from cache effects. Crossing policy (rules vs learned) with placement (cache-aware vs cache-blind) attributes cost differences to decision quality, to cache placement, or to their interaction. A third factor, a model change mid-study, tests durability. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Crossing policy (rules or learned) with placement (cache-aware or cache-blind) separates the cost effect of better decisions from the cost effect of cache placement. If the learned policy wins only in the cache-aware column, its advantage is a placement effect that the rules could be given too. A third factor is a scheduled model change: half the study's duration runs before a planned change of one model in the pool, and half after, without re-fitting. This tests H3.
Analysis plan
Primary analysis. CPVO ratio with a cluster bootstrap interval, resampling task families and then tasks within family. Verified-success difference with a one-sided 95% lower bound. Non-inferiority is assessed by the two-one-sided-tests logic, applied one-sided here (Schuirmann).
Pairing. Where replay provides paired outcomes for the same task under different arms, paired analyses are used. McNemar's test is used for paired binary success.
Multiplicity. H1 is the single confirmatory hypothesis. H2, H3 and secondary metrics are reported with false-discovery-rate control across them (Benjamini & Hochberg) and labeled exploratory.
Verifier error. Verified success depends on verifiers with measured error rates. Verifier versions are pinned for the study, and sensitivity analyses report how conclusions change under the bounds of verifier error. See Evaluating the Evaluators.
Power and sample size
No universal sample size is assumed. The pilot estimates the variance of per-task cost, the baseline success rate and the intra-family correlation, and the confirmatory sample is sized from those. Narrow non-inferiority margins need large samples. [ILLUSTRATIVE EXAMPLE — arithmetic only.] At success rates near 80%, the standard error of a difference between two independent proportions with 500 tasks per arm is about 2.5 percentage points, so a −2-point margin is not demonstrable at that size without pairing. Paired replay and stratification are the main tools for reducing the required sample.
Stopping rules
- Safety stop: any critical incident attributable to the learned policy in Phase 3, such as a hard-constraint violation or an unauthorized effect, halts the online phase pending review.
- Futility: if an interim analysis at a pre-specified point shows that the CPVO interval excludes the 15% threshold in the learned policy's favor with high probability, the study may stop early and keep rules.
- No early stopping for efficacy: a favorable interim result does not end the study early, to avoid inflating apparent gains.
Decision rule
Figure 3 shows the decision rule.
Figure 3. Pre-registered decision rule. The learned policy is promoted only if both pre-registered conditions hold, the effect is not explained by cache placement alone, and no safety stop fired. Thresholds are Ethen proposed targets, fixed before the study. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen proposed targets; not measured.
The learned policy is promoted only if all of the following hold: quality is non-inferior; CPVO is at least 15% lower than strong rules; the decision effect holds in cache-blind cells; and no safety stop fired. Promotion is via canary with rollback. Otherwise, rules remain in production and the full comparison is published, including negative results.
Threats to validity
- Non-stationarity. Prices, provider quality and model versions change during the study. Prices are frozen for accounting, and model versions are pinned where providers allow it.
- Weak comparator. If the rules are weak, the study overstates learning's value. Independent construction and documentation mitigate this.
- Selection bias in logs. Without randomized logging, offline estimates are biased. Phase 1 uses only logged periods with recorded propensities.
- Replay infidelity. Replay estimates are trusted only for task families where replay has been validated against live outcomes.
- Contamination. Development tasks must not leak into evaluation. Family and temporal splits are enforced.
- Verifier drift. Pinned verifier versions, with re-calibration checks at study start and end.
Roles and independence
Four roles are kept separate. The rules team builds and freezes the comparator. The policy team builds the learned policy. The evaluation owner manages sealed task sets, runs the analysis and holds the pre-registration. A reviewer outside both teams adjudicates critical incidents and audits a sample of verified outcomes. The policy team sees development data and aggregate interim results only as the protocol allows. It never sees the confirmatory sample's outcomes before the analysis is locked.
Data and rights
All tasks, logs and replays used in the study must carry rights permitting this research purpose. Tenant-private data stays within its tenant boundary: offline estimation and replay for tenant data run inside that boundary, and only aggregate estimates leave it, as described in Tenant Replay. Randomized exploration in the logging period is bounded by each mandate's permitted configurations and by risk class, and customers are informed where contracts require. Lineage records identify exactly which tasks entered each phase, so that the study can be reproduced or, if rights are withdrawn, its affected results can be identified.
Reporting
The report will include the frozen rules specification, the learned policy's training data lineage and features, all arms' results by family with intervals, the factorial decomposition, the durability result, every critical incident, and every protocol deviation. Data and code are shared to the extent rights permit.
Limitations
This protocol has not been executed. Its thresholds are proposed and may turn out to be too strict or too lenient for particular workloads. The online phase requires live traffic of sufficient volume. Results will apply to the task families studied and the model pool used, and may not generalize.
Conclusion
A learned router should earn its place by beating rules written by competent people, on the outcomes that matter, with quality protected and confounds controlled. This protocol states in advance what evidence would justify promotion, and what evidence would mean the rules stay. Either outcome is worth publishing.
FAQ
Why is frontier-only not the comparator? Because almost any routing policy beats it on cost. Beating strong rules is what shows that learning adds value.
Why randomize at the task level rather than per step? Switching models within a task destroys cache benefits and makes outcomes hard to attribute. Configurations stay fixed within a task.
What happens if the learned policy fails? Rules stay in production and the comparison is published.
Related research
- Why Learned AI Model Routing Must Beat Good Rules — the survey motivating the protocol.
- Faros: Researching How Intelligence Should Choose Intelligence — Faros.
- Why AI Routers Should Log Propensities From Day One — propensity requirement.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI — CPVO definition.
- Evaluating the Evaluators: Reward Integrity for AI Agents — verifier requirement.
References
- Li, H. et al. (2026). LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. arXiv:2601.07206. https://arxiv.org/abs/2601.07206
- Lai, G. et al. (2026). When Routing Collapses: On the Degenerate Convergence of LLM Routers. arXiv:2602.03478. https://arxiv.org/abs/2602.03478
- Mahmood, R. (2026). Routing, Cascades, and User Choice for LLMs. arXiv:2602.09902. https://arxiv.org/abs/2602.09902
- Ding, D. et al. (2024). Hybrid LLM. arXiv:2404.14618. https://arxiv.org/abs/2404.14618
- Ong, I. et al. (2024). RouteLLM. arXiv:2406.18665. https://arxiv.org/abs/2406.18665
- Chen, L., Zaharia, M., Zou, J. (2023). FrugalGPT. arXiv:2305.05176. https://arxiv.org/abs/2305.05176
- Aggarwal, P. et al. (2023). AutoMix. arXiv:2310.12963. https://arxiv.org/abs/2310.12963
- Panda, P. et al. (2025). Adaptive LLM Routing under Budget Constraints. arXiv:2508.21141. https://arxiv.org/abs/2508.21141
- Li, L. et al. (2010). Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms. arXiv:1003.5956. https://arxiv.org/abs/1003.5956
- Dudík, M., Langford, J., Li, L. (2011). Doubly Robust Policy Evaluation and Learning. arXiv:1103.4601. https://arxiv.org/abs/1103.4601
- Yao, S. et al. (2024). τ-bench. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach. J. Pharmacokinet. Biopharm. 15:657–680. https://doi.org/10.1007/BF01068419
- McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2):153–157. https://doi.org/10.1007/BF02295996
- Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
- Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press. https://doi.org/10.1017/9781108653985
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Faros: Researching How Intelligence Should Choose Intelligence
A position paper reframing AI model routing as an execution-configuration decision across model, context, tools, verification, recovery, cost and risk.
- Why Learned AI Model Routing Must Beat Good Rules
A survey of learned LLM routing: what RouteLLM, RouterBench and LLMRouterBench show, why strong rules are the right baseline, and how to test non-inferiority.
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work
A research note on model change assurance: replaying an organization's own historical tasks to find regressions before an LLM upgrade reaches real work.
Explained on the Ethen Blog
- How Ethen Gateway Chooses an Eligible Model
Eligibility is elimination: the Gateway discards every model that cannot serve a request before it ranks anything at all.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.