Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Model Intelligence / Faros

Why AI Routers Should Log Propensities From Day One

Publication type
Research Note
Evidence status
Research Synthesis: Analysis of existing evidence and literature. No new Ethen measurements are implied.technical research note; external literature and Ethen design decisions; no Ethen measurements
Research program
Model Intelligence / Faros
Published
Authors
Ethen Research Lab
Reading time
12 min read

A router that records only what it chose, and not how likely it was to choose it, produces data that cannot answer the question it will most need answered: would a different policy have done better?

Cover image for "Why AI Routers Should Log Propensities From Day One". Decorative abstract motif; contains no data.

Abstract

Every routing decision among models, providers or execution configurations reveals the outcome of one option and hides the outcomes of all the others. Historical routing logs are therefore not a neutral record of how options perform. They are a record of how options performed on the tasks the router chose to send them. This note argues for propensity logging from the first day a router runs: for every decision, record the set of feasible candidates, the probability with which each could have been selected, the action actually taken and the verified outcome. Those four fields are what make off-policy evaluation possible, through inverse propensity weighting, doubly robust estimators and related methods from the contextual-bandit literature. Without them, routing data is confounded by the router's own logic, and attempts to reconstruct propensities after the fact are weak. We describe a minimal logging schema, explain how a small, safety-bounded amount of randomization creates the overlap these estimators require, and set out what propensity logging does not fix: model-version drift, cache interference, verifier error and the arrival of options that were never logged. Propensity logging does not by itself solve causal inference. It is the inexpensive precondition without which most later questions about routing cannot be answered credibly.

One outcome per decision

A router sends a task to one configuration. The organization learns whether that configuration succeeded, what it cost and how long it took. It does not learn what the other configurations would have done on the same task. The contextual-bandit literature calls this partial or bandit feedback, and it is the defining difficulty of evaluating decision policies from logs (Li et al., 2010a; Dudík et al.).

The difficulty is easy to underestimate because routing logs are large. A router that has made a million decisions appears to have a million data points about model performance. It has a million data points about the performance of chosen models on the tasks for which they were chosen. If the router sends hard tasks to the strongest model, the strongest model's logged success rate understates its ability and the weaker model's overstates its own. A new policy trained or evaluated naively on those logs will learn the router's habits, not the models' capabilities.

The broader decision layer in which this matters, choosing models, tools, context, verification and recovery for each task, is described in Faros. The case that learned routing must earn its place against strong rules is made in Why Learned AI Model Routing Must Beat Good Rules. Both depend on being able to estimate how a policy would have performed without deploying it. That estimate depends on what was logged.

How routing logs become selection-biased

Figure 1 shows the core problem as a causal diagram.

Causal diagram. Task features (family, length, difficulty, tools, tier) point to both router choice and verified outcome. Router choice points to verified outcome. Unlogged features, such as signals the router read from the raw prompt, also point to both choice and outcome. Cache state and model version drift point to the outcome. A separate box labeled logged randomization with recorded propensity points into router choice, breaking the dependence of choice on unobserved features within the explored share.

Figure 1. Why routing logs are confounded. Task features influence both which configuration the router chooses and whether the task succeeds. Without logged selection probabilities, the effect of the configuration cannot be separated from the router's own sorting of tasks. Bounded randomization with recorded propensities restores the missing variation. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen research note; standard causal reasoning for logged bandit feedback.

Task features, such as length, family, apparent difficulty, required tools and customer tier, influence both the router's choice and the outcome. That makes them confounders. Some of these features are logged; some are not, including whatever the router inferred from the raw prompt. Four further mechanisms make the bias worse in practice.

Task-family confounding. Rules often route whole families to particular configurations. If code tasks always go to one model and support tasks to another, the logs contain no information about how either model performs on the other family. No amount of statistical adjustment can recover a comparison that the data never contained.

Deterministic policies. A rule-based router assigns each context to one action with probability one. Its logs contain no variation within a context, so the support required by off-policy estimators is absent. The problem is not bias that can be corrected but missing information.

Cache effects. The cost and latency of a decision depend on whether the chosen provider has a warm cache for the task's prefix, which depends on earlier decisions. A configuration may look cheap because the router kept sending related tasks to it. This is interference between decisions, which standard estimators assume away.

Non-stationarity. Model versions, prices and provider health change during the logging period. A logged action called "model X" in March may not be the same action in June, even if the identifier did not change (Chen et al.).

The four fields, and what surrounds them

Figure 2 shows a minimal logging schema.

Table of log fields with purpose and example content. Accented essential fields: candidate set (feasible configurations after hard constraints); selection probability per candidate; chosen action with exact model, provider and version; verified outcome with provisional flag, cost and latency. Supporting fields: policy version and features used; cache status at decision time; timestamp and price table version; rights record of the task.

Figure 2. A minimal routing log record. The four essential fields, in the accented rows, enable off-policy evaluation. The surrounding fields make each record interpretable after policies, versions and prices have changed. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research note; Ethen design decision on propensity logging.

The four essential fields are:

  • Candidate set. Every configuration that was feasible for this task under the governing mandate and tenant policy, after hard constraints such as residency, budget and capability were applied. Configurations that were forbidden are not candidates and must never be explored.
  • Selection probability. The probability with which the logging policy would have chosen each candidate, or at minimum the chosen one. For a deterministic rule this is one for the chosen action and zero for the rest. Recording that is still useful: it documents that there is no overlap.
  • Chosen action. The configuration actually executed, with exact model and provider identifiers, any version metadata the provider exposes, and the tool, skill and verifier versions in force.
  • Observed outcome. The verified outcome, not merely whether the run completed, together with cost, latency and escalation. Outcomes that are still provisional, because a reopen or revert window has not closed, are marked as such and updated.

Around these sit the fields that make the record interpretable later: the policy version that made the decision, the features it used, the cache status at decision time, a timestamp and the rights record of the task. In the Ethen architecture, the routing record lives inside the work receipt for the task, so that outcome, cost and decision can be joined without a separate pipeline.

What the fields make possible

With logged propensities and overlap, several established estimators become available.

Inverse propensity weighting. The value of a new policy is estimated by reweighting logged outcomes by the ratio of the new policy's probability of each action to the logging policy's probability. The idea goes back to the Horvitz–Thompson estimator in survey sampling (Horvitz & Thompson). It is unbiased under its assumptions but can have high variance when logging probabilities are small.

Replay evaluation. When the logging policy chose uniformly at random, a new policy can be evaluated by keeping only the logged decisions where it agrees with the logged action. Li et al. showed this replay method provides unbiased offline evaluation of contextual-bandit algorithms and compared it with online bucket tests on news recommendation data (Li et al., 2010a).

Doubly robust estimation. Reward models alone tend to be biased, and propensity weighting alone tends to be noisy. Doubly robust estimators combine the two and yield accurate value estimates when either the reward model or the model of the past policy is good (Dudík et al.).

Learning from logged feedback. The same weights support training new policies, not only evaluating them. Counterfactual risk minimization adds a variance penalty to the propensity-weighted objective so that the learned policy does not exploit the estimator's noise (Swaminathan & Joachims).

These methods are mature in advertising and recommendation, where large systems have used causal reasoning on logged data to predict the effect of changes (Bottou et al.), and the first contextual-bandit treatments of news recommendation established the logging discipline they require (Li et al., 2010b). Routing among models is a close analogue. Bandit approaches have already been applied to budgeted model routing (Panda et al.). What is often missing in practice is not the estimator but the logged data it needs.

Exploration without unsafe exploration

Overlap requires that the logging policy sometimes chooses options other than its favorite. In a production agent system, that raises an obvious concern: randomly sending work to a weaker configuration can cause failures.

The answer is to bound exploration rather than avoid it. Randomization happens only within the feasible set, so it never selects a configuration the mandate forbids. It is restricted to low-risk task classes, and the probability mass assigned to non-default options is small and recorded. High-risk classes, such as tasks with irreversible effects, can be excluded entirely, at the price of having no off-policy evidence for them; for those, paired replay in a stubbed environment, described in Counterfactual Replay for AI Agents, is the safer source of counterfactual evidence. A permanent small holdout on the incumbent policy, a standard practice in online experimentation (Kohavi et al.), also continues to measure the counterfactual after any new policy is deployed.

Randomization should be at the task level with stickiness within a task. Switching configurations mid-task confounds attribution and destroys cache benefits.

Why reconstruction after the fact is weak

Organizations that did not log propensities sometimes try to estimate them later by fitting a model that predicts which action the router chose from the logged features. This is the standard approach in observational studies, and it is much weaker here than it looks.

Figure 3 summarizes the decision.

Decision tree. Question 1: were selection probabilities logged at decision time? If no, reconstructed propensities support exploratory analysis only; use replay or a randomized experiment for decisions. If yes, question 2: does the new policy choose only actions that had non-zero logged probability in those contexts? If no, the unsupported region needs replay or new exploration. If yes, question 3: is the effective sample size adequate after clipping? If no, collect more bounded exploration. If yes, report IPS and doubly robust estimates with intervals.

Figure 3. Can a new routing policy be evaluated from the logs?. Off-policy evaluation is credible only when propensities were logged and the new policy's choices fall within the logged support. Otherwise the honest options are replay, a randomized experiment, or an exploratory analysis labeled as such. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen research note; off-policy evaluation literature.

First, a rule-based router is close to deterministic, so the fitted propensities will be near zero or one, and the weights explode or the effective sample size collapses. Second, the router often used features that were not logged, including the raw prompt, so the fitted model cannot recover the true assignment mechanism, and unobserved confounding remains. Third, routing rules change over time, often without versioned records, so a single fitted model averages over several policies. Fourth, any reconstruction relies on assumptions that cannot be checked from the same data. Reconstructed propensities can support exploratory analysis. They should not support a decision to replace a production policy.

What propensity logging does not fix

Propensity logging is necessary for credible off-policy evaluation of routing. It is not sufficient.

  • New options. A model released after the logging period was never a candidate, so no logged weight can evaluate a policy that uses it. That requires replay or a fresh experiment.
  • Version drift. If the action called "model X" changed behavior during logging, pooled estimates mix two different actions. Logging version metadata and splitting by period mitigate this; they do not remove it.
  • Interference. Cache effects and shared rate limits mean one decision can change another's outcome. Estimators that assume independent decisions will misattribute some cost differences.
  • Outcome quality. The estimates are only as good as the verified outcomes they weight. Verifier error passes straight through into policy evaluation.
  • Variance. With small exploration rates, estimates for policies that differ greatly from the logging policy will be noisy. Effective sample size should be reported with every estimate, and weights clipped with the bias that clipping introduces stated.

The protocol that uses these logs alongside paired replay and a randomized online comparison is set out in How to Test Whether Learned AI Routing Beats Strong Rules.

A note on routing benchmarks

Public routing benchmarks evaluate routers on datasets where every model's response to every query is available, which removes the partial-feedback problem by construction (Hu et al.; Li et al., 2026). A large-scale re-evaluation using such a benchmark found that many routing methods perform similarly and several fail to reliably outperform a simple baseline (Li et al., 2026), and routers trained to predict scalar scores can collapse onto the most expensive model (Lai et al.). Full-information benchmarks are valuable for method development. Production routing never has full information, which is exactly why the logging decision matters.

Limitations

This is a technical note grounded in external literature and Ethen design decisions. It reports no Ethen measurements of logging overhead, exploration cost or estimator accuracy on agent workloads. The appropriate exploration rate, and which task classes can tolerate it, must be determined per deployment and may be small enough that estimates for some policies remain imprecise. Agent tasks are longer and more heterogeneous than the recommendation settings in which these estimators were developed, and their behavior there is an empirical question.

Conclusion

Logging propensities costs a few fields per decision. Not logging them makes a router's history nearly useless for the question that matters when anyone proposes a better policy. A router should record what was possible, how likely each choice was, what was done and what was verified, from the first decision it makes, because that data cannot be recreated afterwards.

FAQ

Does a rule-based router need propensity logging? Yes. Logging the candidate set and a probability of one documents where there is no overlap, and adding small bounded randomization in low-risk classes creates the evidence that later evaluation will need.

Can we estimate propensities later from the logs? Only weakly. Near-deterministic rules, unlogged features and unversioned rule changes make reconstructed propensities unreliable for decisions.

Does this replace replay or online experiments? No. It complements them. Logged data cannot evaluate options that were never candidates, and high-risk classes may never be explored.

References

  1. Li, L. et al. (2010a). Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms. arXiv:1003.5956. https://arxiv.org/abs/1003.5956
  2. Dudík, M., Langford, J., Li, L. (2011). Doubly Robust Policy Evaluation and Learning. arXiv:1103.4601. https://arxiv.org/abs/1103.4601
  3. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  4. Horvitz, D. G., Thompson, D. J. (1952). A Generalization of Sampling Without Replacement From a Finite Universe. JASA 47(260). https://doi.org/10.1080/01621459.1952.10483446
  5. Swaminathan, A., Joachims, T. (2015). Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. arXiv:1502.02362. https://arxiv.org/abs/1502.02362
  6. Bottou, L. et al. (2012). Counterfactual Reasoning and Learning Systems. arXiv:1209.2355. https://arxiv.org/abs/1209.2355
  7. Li, L. et al. (2010b). A Contextual-Bandit Approach to Personalized News Article Recommendation. arXiv:1003.0146. https://arxiv.org/abs/1003.0146
  8. Panda, P. et al. (2025). Adaptive LLM Routing under Budget Constraints. arXiv:2508.21141. https://arxiv.org/abs/2508.21141
  9. Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press. https://doi.org/10.1017/9781108653985
  10. Hu, Q. J. et al. (2024). RouterBench: A Benchmark for Multi-LLM Routing System. arXiv:2403.12031. https://arxiv.org/abs/2403.12031
  11. Li, H. et al. (2026). LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. arXiv:2601.07206. https://arxiv.org/abs/2601.07206
  12. Lai, G. et al. (2026). When Routing Collapses: On the Degenerate Convergence of LLM Routers. arXiv:2602.03478. https://arxiv.org/abs/2602.03478

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.