Research Note · 2026-10-03 · Model Intelligence / Faros
Why AI Routers Should Log Propensities From Day One
A router that records only what it chose, and not how likely it was to choose it, produces data that cannot answer the question it will most need answered: would a different policy have done better?
Abstract
Every routing decision among models, providers or execution configurations reveals the outcome of one option and hides the outcomes of all the others. Historical routing logs are therefore not a neutral record of how options perform. They are a record of how options performed on the tasks the router chose to send them. This note argues for propensity logging from the first day a router runs: for every decision, record the set of feasible candidates, the probability with which each could have been selected, the action actually taken and the verified outcome. Those four fields are what make off-policy evaluation possible, through inverse propensity weighting, doubly robust estimators and related methods from the contextual-bandit literature. Without them, routing data is confounded by the router's own logic, and attempts to reconstruct propensities after the fact are weak. We describe a minimal logging schema, explain how a small, safety-bounded amount of randomization creates the overlap these estimators require, and set out what propensity logging does not fix: model-version drift, cache interference, verifier error and the arrival of options that were never logged. Propensity logging does not by itself solve causal inference. It is the inexpensive precondition without which most later questions about routing cannot be answered credibly.
One outcome per decision
A router sends a task to one configuration. The organization learns whether that configuration succeeded, what it cost and how long it took. It does not learn what the other configurations would have done on the same task. The contextual-bandit literature calls this partial or bandit feedback, and it is the defining difficulty of evaluating decision policies from logs (Li et al., 2010a; Dudík et al.).
The difficulty is easy to underestimate because routing logs are large. A router that has made a million decisions appears to have a million data points about model performance. It has a million data points about the performance of chosen models on the tasks for which they were chosen. If the router sends hard tasks to the strongest model, the strongest model's logged success rate understates its ability and the weaker model's overstates its own. A new policy trained or evaluated naively on those logs will learn the router's habits, not the models' capabilities.
The broader decision layer in which this matters, choosing models, tools, context, verification and recovery for each task, is described in Faros. The case that learned routing must earn its place against strong rules is made in Why Learned AI Model Routing Must Beat Good Rules. Both depend on being able to estimate how a policy would have performed without deploying it. That estimate depends on what was logged.
How routing logs become selection-biased
Figure 1 shows the core problem as a causal diagram.
Figure 1. Why routing logs are confounded. Task features influence both which configuration the router chooses and whether the task succeeds. Without logged selection probabilities, the effect of the configuration cannot be separated from the router's own sorting of tasks. Bounded randomization with recorded propensities restores the missing variation. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen research note; standard causal reasoning for logged bandit feedback.
Task features, such as length, family, apparent difficulty, required tools and customer tier, influence both the router's choice and the outcome. That makes them confounders. Some of these features are logged; some are not, including whatever the router inferred from the raw prompt. Four further mechanisms make the bias worse in practice.
Task-family confounding. Rules often route whole families to particular configurations. If code tasks always go to one model and support tasks to another, the logs contain no information about how either model performs on the other family. No amount of statistical adjustment can recover a comparison that the data never contained.
Deterministic policies. A rule-based router assigns each context to one action with probability one. Its logs contain no variation within a context, so the support required by off-policy estimators is absent. The problem is not bias that can be corrected but missing information.
Cache effects. The cost and latency of a decision depend on whether the chosen provider has a warm cache for the task's prefix, which depends on earlier decisions. A configuration may look cheap because the router kept sending related tasks to it. This is interference between decisions, which standard estimators assume away.
Non-stationarity. Model versions, prices and provider health change during the logging period. A logged action called "model X" in March may not be the same action in June, even if the identifier did not change (Chen et al.).
The four fields, and what surrounds them
Figure 2 shows a minimal logging schema.
Figure 2. A minimal routing log record. The four essential fields, in the accented rows, enable off-policy evaluation. The surrounding fields make each record interpretable after policies, versions and prices have changed. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research note; Ethen design decision on propensity logging.
The four essential fields are:
- Candidate set. Every configuration that was feasible for this task under the governing mandate and tenant policy, after hard constraints such as residency, budget and capability were applied. Configurations that were forbidden are not candidates and must never be explored.
- Selection probability. The probability with which the logging policy would have chosen each candidate, or at minimum the chosen one. For a deterministic rule this is one for the chosen action and zero for the rest. Recording that is still useful: it documents that there is no overlap.
- Chosen action. The configuration actually executed, with exact model and provider identifiers, any version metadata the provider exposes, and the tool, skill and verifier versions in force.
- Observed outcome. The verified outcome, not merely whether the run completed, together with cost, latency and escalation. Outcomes that are still provisional, because a reopen or revert window has not closed, are marked as such and updated.
Around these sit the fields that make the record interpretable later: the policy version that made the decision, the features it used, the cache status at decision time, a timestamp and the rights record of the task. In the Ethen architecture, the routing record lives inside the work receipt for the task, so that outcome, cost and decision can be joined without a separate pipeline.
What the fields make possible
With logged propensities and overlap, several established estimators become available.
Inverse propensity weighting. The value of a new policy is estimated by reweighting logged outcomes by the ratio of the new policy's probability of each action to the logging policy's probability. The idea goes back to the Horvitz–Thompson estimator in survey sampling (Horvitz & Thompson). It is unbiased under its assumptions but can have high variance when logging probabilities are small.
Replay evaluation. When the logging policy chose uniformly at random, a new policy can be evaluated by keeping only the logged decisions where it agrees with the logged action. Li et al. showed this replay method provides unbiased offline evaluation of contextual-bandit algorithms and compared it with online bucket tests on news recommendation data (Li et al., 2010a).
Doubly robust estimation. Reward models alone tend to be biased, and propensity weighting alone tends to be noisy. Doubly robust estimators combine the two and yield accurate value estimates when either the reward model or the model of the past policy is good (Dudík et al.).
Learning from logged feedback. The same weights support training new policies, not only evaluating them. Counterfactual risk minimization adds a variance penalty to the propensity-weighted objective so that the learned policy does not exploit the estimator's noise (Swaminathan & Joachims).
These methods are mature in advertising and recommendation, where large systems have used causal reasoning on logged data to predict the effect of changes (Bottou et al.), and the first contextual-bandit treatments of news recommendation established the logging discipline they require (Li et al., 2010b). Routing among models is a close analogue. Bandit approaches have already been applied to budgeted model routing (Panda et al.). What is often missing in practice is not the estimator but the logged data it needs.
Exploration without unsafe exploration
Overlap requires that the logging policy sometimes chooses options other than its favorite. In a production agent system, that raises an obvious concern: randomly sending work to a weaker configuration can cause failures.
The answer is to bound exploration rather than avoid it. Randomization happens only within the feasible set, so it never selects a configuration the mandate forbids. It is restricted to low-risk task classes, and the probability mass assigned to non-default options is small and recorded. High-risk classes, such as tasks with irreversible effects, can be excluded entirely, at the price of having no off-policy evidence for them; for those, paired replay in a stubbed environment, described in Counterfactual Replay for AI Agents, is the safer source of counterfactual evidence. A permanent small holdout on the incumbent policy, a standard practice in online experimentation (Kohavi et al.), also continues to measure the counterfactual after any new policy is deployed.
Randomization should be at the task level with stickiness within a task. Switching configurations mid-task confounds attribution and destroys cache benefits.
Why reconstruction after the fact is weak
Organizations that did not log propensities sometimes try to estimate them later by fitting a model that predicts which action the router chose from the logged features. This is the standard approach in observational studies, and it is much weaker here than it looks.
Figure 3 summarizes the decision.
Figure 3. Can a new routing policy be evaluated from the logs?. Off-policy evaluation is credible only when propensities were logged and the new policy's choices fall within the logged support. Otherwise the honest options are replay, a randomized experiment, or an exploratory analysis labeled as such. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen research note; off-policy evaluation literature.
First, a rule-based router is close to deterministic, so the fitted propensities will be near zero or one, and the weights explode or the effective sample size collapses. Second, the router often used features that were not logged, including the raw prompt, so the fitted model cannot recover the true assignment mechanism, and unobserved confounding remains. Third, routing rules change over time, often without versioned records, so a single fitted model averages over several policies. Fourth, any reconstruction relies on assumptions that cannot be checked from the same data. Reconstructed propensities can support exploratory analysis. They should not support a decision to replace a production policy.
What propensity logging does not fix
Propensity logging is necessary for credible off-policy evaluation of routing. It is not sufficient.
- New options. A model released after the logging period was never a candidate, so no logged weight can evaluate a policy that uses it. That requires replay or a fresh experiment.
- Version drift. If the action called "model X" changed behavior during logging, pooled estimates mix two different actions. Logging version metadata and splitting by period mitigate this; they do not remove it.
- Interference. Cache effects and shared rate limits mean one decision can change another's outcome. Estimators that assume independent decisions will misattribute some cost differences.
- Outcome quality. The estimates are only as good as the verified outcomes they weight. Verifier error passes straight through into policy evaluation.
- Variance. With small exploration rates, estimates for policies that differ greatly from the logging policy will be noisy. Effective sample size should be reported with every estimate, and weights clipped with the bias that clipping introduces stated.
The protocol that uses these logs alongside paired replay and a randomized online comparison is set out in How to Test Whether Learned AI Routing Beats Strong Rules.
A note on routing benchmarks
Public routing benchmarks evaluate routers on datasets where every model's response to every query is available, which removes the partial-feedback problem by construction (Hu et al.; Li et al., 2026). A large-scale re-evaluation using such a benchmark found that many routing methods perform similarly and several fail to reliably outperform a simple baseline (Li et al., 2026), and routers trained to predict scalar scores can collapse onto the most expensive model (Lai et al.). Full-information benchmarks are valuable for method development. Production routing never has full information, which is exactly why the logging decision matters.
Limitations
This is a technical note grounded in external literature and Ethen design decisions. It reports no Ethen measurements of logging overhead, exploration cost or estimator accuracy on agent workloads. The appropriate exploration rate, and which task classes can tolerate it, must be determined per deployment and may be small enough that estimates for some policies remain imprecise. Agent tasks are longer and more heterogeneous than the recommendation settings in which these estimators were developed, and their behavior there is an empirical question.
Conclusion
Logging propensities costs a few fields per decision. Not logging them makes a router's history nearly useless for the question that matters when anyone proposes a better policy. A router should record what was possible, how likely each choice was, what was done and what was verified, from the first decision it makes, because that data cannot be recreated afterwards.
FAQ
Does a rule-based router need propensity logging? Yes. Logging the candidate set and a probability of one documents where there is no overlap, and adding small bounded randomization in low-risk classes creates the evidence that later evaluation will need.
Can we estimate propensities later from the logs? Only weakly. Near-deterministic rules, unlogged features and unversioned rule changes make reconstructed propensities unreliable for decisions.
Does this replace replay or online experiments? No. It complements them. Logged data cannot evaluate options that were never candidates, and high-risk classes may never be explored.
Related research
- Faros: Researching How Intelligence Should Choose Intelligence — Faros.
- Counterfactual Replay for AI Agents — replay complements OPE.
- How to Test Whether Learned AI Routing Beats Strong Rules — the routing protocol.
- Why Learned AI Model Routing Must Beat Good Rules — learned routing literature.
References
- Li, L. et al. (2010a). Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms. arXiv:1003.5956. https://arxiv.org/abs/1003.5956
- Dudík, M., Langford, J., Li, L. (2011). Doubly Robust Policy Evaluation and Learning. arXiv:1103.4601. https://arxiv.org/abs/1103.4601
- Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
- Horvitz, D. G., Thompson, D. J. (1952). A Generalization of Sampling Without Replacement From a Finite Universe. JASA 47(260). https://doi.org/10.1080/01621459.1952.10483446
- Swaminathan, A., Joachims, T. (2015). Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. arXiv:1502.02362. https://arxiv.org/abs/1502.02362
- Bottou, L. et al. (2012). Counterfactual Reasoning and Learning Systems. arXiv:1209.2355. https://arxiv.org/abs/1209.2355
- Li, L. et al. (2010b). A Contextual-Bandit Approach to Personalized News Article Recommendation. arXiv:1003.0146. https://arxiv.org/abs/1003.0146
- Panda, P. et al. (2025). Adaptive LLM Routing under Budget Constraints. arXiv:2508.21141. https://arxiv.org/abs/2508.21141
- Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press. https://doi.org/10.1017/9781108653985
- Hu, Q. J. et al. (2024). RouterBench: A Benchmark for Multi-LLM Routing System. arXiv:2403.12031. https://arxiv.org/abs/2403.12031
- Li, H. et al. (2026). LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. arXiv:2601.07206. https://arxiv.org/abs/2601.07206
- Lai, G. et al. (2026). When Routing Collapses: On the Degenerate Convergence of LLM Routers. arXiv:2602.03478. https://arxiv.org/abs/2602.03478
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Faros: Researching How Intelligence Should Choose Intelligence
A position paper reframing AI model routing as an execution-configuration decision across model, context, tools, verification, recovery, cost and risk.
- Why Learned AI Model Routing Must Beat Good Rules
A survey of learned LLM routing: what RouteLLM, RouterBench and LLMRouterBench show, why strong rules are the right baseline, and how to test non-inferiority.
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work
A research note on model change assurance: replaying an organization's own historical tasks to find regressions before an LLM upgrade reaches real work.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.