Skip to content

EthenEthenEthen

Research Proposal · 2026-10-03 · Adaptive Intelligence

Counterfactual Replay for AI Agents

Publication type
Research Proposal
Research program
Adaptive Intelligence
Published
Authors
Ethen Research Lab
Reading time
12 min read

Every completed task leaves a question behind: would a different model, tool, context strategy or recovery policy have done better? Replay is how we propose to answer it.

Cover image for "Counterfactual Replay for AI Agents". Decorative abstract motif; contains no data.

Abstract

Production logs show what an agent did under the configuration it actually used. They cannot show what would have happened under another one. This proposal describes counterfactual replay: capturing completed tasks as replayable snapshots and re-running them, with side effects stubbed, under alternative models, providers, tool sets, verifier policies, context strategies and recovery strategies, then scoring every run with the same verifiers. The product of replay is a table of verified outcomes per task and configuration, the data that routing, model-upgrade assurance, recovery learning and context research all need. We explain why replay complements, rather than replaces, off-policy evaluation from logged exploration. We describe four replay modes and their fidelity trade-offs, and we identify the central risk: replay results that diverge from what would happen live. Before replay can drive decisions, its agreement with live outcomes must be measured, and we propose a validation design for that. Counterfactual evaluation of AI agents in this sense is an Ethen research proposal; no replay system has been built or validated.

The counterfactual problem

Every decision an agent system makes, which model, which tools, how much context, whether to verify, how to recover, is evaluated only along the path actually taken. If the system used model X and the task succeeded, we do not know whether model Y would have succeeded more cheaply. If it retried after a timeout and created a duplicate, we do not know whether reconciling first would have avoided it. Learning from logs in this situation is learning from bandit feedback: one arm observed per round, with outcomes for other arms missing. That missing data is the central obstacle to improving such systems (Bottou et al.).

There are two established ways to recover counterfactual information.

Randomized logging with known propensities. If the system sometimes chooses among feasible options at random, with recorded probabilities, then the value of any other policy can be estimated without bias from the logs, using inverse propensity scoring or doubly robust estimators (Li et al., 2010; Dudík et al.), and new policies can be learned from the same logs (Swaminathan & Joachims). This works at scale and on real traffic, but it requires exploration, which has costs, and estimates have high variance when the new policy differs greatly from the logging policy. See Why AI Routers Should Log Propensities From Day One.

Re-execution. If the task can be run again from the same starting point under a different configuration, the counterfactual can be observed directly. This is what replay offers. It provides paired comparisons on identical tasks, which are far more statistically efficient than unpaired ones, for the same reason that variance reduction matters in online controlled experiments (Kohavi et al.). It requires no exploration on live traffic. Its cost is environmental fidelity: replay is only as informative as the replayed environment is faithful.

The two approaches are complementary. Logged exploration measures what happens on real traffic but estimates only noisily for policies far from the logging policy. Replay measures directly on identical tasks but in an environment that may not behave like production.

What replay produces

Figure 1 shows the pipeline.

Pipeline: completed task and its receipt; snapshot (initial state, spec, recorded tool responses, user turns, verifier definitions); replay harness fans out into configurations A through D varying model, tools, context strategy, recovery strategy; each produces a verified outcome with cost; the results form a row in a counterfactual outcome table keyed by task and configuration.

Figure 1. Replaying one completed task under alternative configurations. A completed task is captured as a replayable snapshot: initial state, task specification, recorded tool responses and user turns. The replay harness runs the task under each alternative configuration with side effects stubbed, and the same verifiers score every run. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research proposal.

A completed task, with its Work Receipt, is converted into a replayable snapshot: the initial state of relevant systems, the task specification, the recorded responses of tools the agent called, the recorded user turns, and the verifier definitions used to judge the outcome. A replay harness then runs the task under each alternative configuration. Side-effecting tools are stubbed, and every run is scored with the same verifiers. The result is a row of a counterfactual outcome table: for one task, the verified outcome and cost under each configuration.

A table of this kind, accumulated over many tasks, supports several research programs:

  • Decision policy. Which configuration produces verified outcomes most cheaply for which kinds of task: the core question for Faros.
  • Model change. Which tasks regress, and why, if a tenant switches from one model to another: Model Change Assurance.
  • Recovery. Which recovery strategy succeeds from a given failure state: the branching mechanism of the Recovery Atlas.
  • Context. How much context a task needs, by replaying it under different context budgets and compaction strategies.

An illustrative walkthrough

[ILLUSTRATIVE EXAMPLE — not an Ethen result.] Suppose an agent resolved a billing dispute using a large model. It read the ticket, queried invoices, found a duplicate charge, proposed a refund, obtained approval and issued it. The snapshot captures the ticket, the invoice records, the payment-system state and the approval rule. Three alternatives are replayed. A smaller model reads the same ticket, queries the same invoices and proposes the same refund; replay serves the recorded responses and the verifier confirms the correct refund was proposed at a fraction of the cost. A second configuration exposes fewer tools; the agent cannot find the duplicate charge and escalates, which the verifier scores as a safe but incomplete outcome. A third uses a different recovery policy after an injected payment timeout; it reconciles before retrying and avoids a duplicate refund that the original policy would have issued. One completed task has produced three verified counterfactuals, two of them informative precisely because the alternative behaved differently from the original.

The divergence problem

Replay is straightforward when the alternative configuration takes exactly the same actions as the original run. It is hard, and most informative, when the alternative takes different actions. If the original agent queried a database for invoices and the alternative agent instead queries for payments, there is no recorded response for the alternative query. The replay must decide how to respond. Figure 2 compares four ways of doing so.

Matrix comparing four replay modes (record-and-replay of tool responses, emulated environment, reconstructed sandbox, live shadow without effects) on: supports divergent actions, side-effect safety, fidelity to real system, cost per replay. Record-and-replay is safe and cheap but cannot handle divergent actions; reconstructed sandboxes handle divergence with high fidelity at higher cost; emulation handles divergence at uncertain fidelity.

Figure 2. Replay modes and what they can support. Replay fidelity depends on how the environment responds when an alternative configuration takes actions the original run never took. No single mode is both faithful and safe for every task; choosing the mode is part of the experimental design. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative.

Record-and-replay returns recorded responses for actions the original run took and fails, or ends the replay, for anything else. It is cheap and safe, and faithful exactly where it applies. It cannot evaluate configurations that behave differently, which are the interesting ones.

Emulated environments use a model to simulate responses to new actions. LM-emulated sandboxes make it cheap to explore risky actions without executing them (Ruan et al.). The price is uncertain fidelity: the emulator may respond as a real system would not, and an agent evaluated against an emulator may be learning the emulator's quirks.

Reconstructed sandboxes restore the relevant state into real or high-fidelity replica systems, such as containers with a copy of a repository, a database restored from a snapshot, or a replica of a ticketing system. Divergent actions then receive real responses from real software. This is the mode used by benchmarks that grade agents on environment state (Jimenez et al.; Trivedi et al.; Drouin et al.). Fidelity is high if the relevant state was captured; cost is higher.

Live shadow, read-only runs the alternative configuration against live systems with all writes blocked. Reads are real; writes are stubbed. It is faithful for read-heavy tasks and useless for tasks whose outcome depends on writes.

In practice, a replay system will mix modes by task family. Code tasks can often be reconstructed in containers with high fidelity. Workflows across several SaaS systems may need emulation or replicas, as in a synthetic enterprise.

What cannot be replayed safely

Some effects cannot be stubbed without losing the point of the evaluation, and some must never be repeated. Sending a real email, moving real money, or modifying a production system cannot be part of replay. Replay stubs these effects and judges the proposed action rather than its consequences. That is a weaker form of evaluation, and replay results for effect-heavy tasks should be labeled accordingly. The note on unknown effects explains why ambiguous effects are especially hazardous. A replay system that accidentally executes a live side effect is a serious incident, so stubbing must be enforced by the environment, not by the agent's instructions.

Validating replay against live outcomes

The central risk of replay is that its results diverge from what would happen live. If replay says model Y would have succeeded on 85% of tasks where model X succeeded on 80%, that comparison is meaningful only if replay faithfully predicts live outcomes for both. We therefore propose that replay be validated before its results drive any decision (Figure 3).

Experiment flow: sample tasks with consent; run configuration K live (or in its original logged run); replay the same tasks under configuration K; compare verified outcomes pairwise; compute agreement and divergence by family and replay mode; if agreement is below threshold, restrict replay to families where it passes.

Figure 3. Validating replay against live outcomes. Before replay results drive decisions, their agreement with live outcomes must be measured. A sample of tasks is executed live under a configuration and also replayed under that same configuration; disagreement rates, stratified by task family and replay mode, bound how far replay can be trusted. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Sample rights-cleared tasks. For a configuration K, compare each task's live or originally logged outcome with its replayed outcome under the same configuration K. Agreement between the two, stratified by task family and replay mode, measures how far replay can be trusted. Where agreement falls below a pre-registered threshold, replay results for that family should not be used for decisions until the environment improves. Disagreements should be examined individually: they reveal which parts of the environment the snapshot failed to capture.

Two subtler checks are needed. Stochasticity: agents are not deterministic, so replaying the same configuration twice can produce different outcomes. Measuring replay-vs-replay agreement first establishes the noise floor that replay-vs-live agreement must be compared against. The pass^k reliability metric of τ-bench captures a related idea: the probability that an agent succeeds on all of k independent trials (Yao et al.). Temporal drift: a snapshot captures the world at one time. Replaying a month later against a model whose provider has changed it silently may measure the model change rather than the configuration. Hosted model behavior is documented to shift between versions (Chen et al.), so replays must pin model versions where providers allow it and record them where they do not.

Privacy and boundaries

Replay snapshots contain customer state and recorded responses that are often sensitive. Under the default rights model, snapshots are tenant-private and replay runs inside the tenant's boundary. Only aggregate results, such as agreement rates, outcome deltas by configuration and cost differences, may leave, and only where a purpose grant permits. The deployment architecture for in-boundary replay is described in Tenant Replay.

Research questions

  1. Fidelity. For which task families does replay agree with live outcomes at a useful level, and in which replay modes?
  2. Efficiency. How much does pairing on identical tasks reduce the number of tasks needed to detect a given difference between configurations, compared with unpaired live comparisons?
  3. Complementarity. Do off-policy estimates from logged exploration and replay estimates agree where both are available? Where they disagree, which is closer to subsequent live results?
  4. Cost. What does a replay cost relative to the decision value it informs? A replay infrastructure that costs more than the decisions it improves is not worth building.

Failure modes

  • Emulator overfitting. Configurations tuned against an emulated environment may exploit its quirks. Results from emulated replay should be confirmed in higher-fidelity modes before they shape production.
  • Snapshot incompleteness. A task may depend on state that was never captured: a cache, a rate limit, a time-dependent API. Replay will then differ from live in ways that look like configuration effects.
  • Verifier coupling. If the verifier itself relies on recorded responses, it may judge divergent runs unfairly. Verifiers for replay must be defined on state, not on matching the original trajectory.
  • Survivorship. Replaying only completed tasks excludes tasks that were abandoned or crashed before a snapshot was captured, biasing the corpus toward easier work.

Limitations

No replay system has been built or validated for this proposal. Fidelity for multi-system enterprise workflows is unknown and may be poor. Costs could make replay viable only for a sample of tasks. The counterfactual table is only as good as the verifiers that score it, so every limitation discussed in Evaluating the Evaluators applies here.

Conclusion

Logs tell us what happened; replay can tell us what would have happened. For agent systems that must choose among models, tools, context strategies and recoveries, that counterfactual is the missing data. Replay can supply it directly and with paired precision, but only to the degree that the replayed environment behaves like the real one. The first deliverable of a replay program is therefore not a leaderboard of configurations. It is a measured map of where replay can be trusted.

FAQ

How is counterfactual replay different from A/B testing? A/B testing assigns live tasks to configurations at random and observes each task under one configuration. Replay runs the same task under several configurations in a controlled environment, so differences are paired, but the environment may not match live behavior.

Does replay execute real side effects? No. Side-effecting tools are stubbed by the environment. Tasks whose outcome depends on real effects are evaluated on the proposed actions, which is a weaker form of evaluation.

Does replay replace off-policy evaluation? No. They complement each other: off-policy evaluation uses real traffic but is noisy for policies far from the logging policy; replay is precise but depends on environment fidelity.

References

  1. Bottou, L. et al. (2012). Counterfactual Reasoning and Learning Systems. arXiv:1209.2355. https://arxiv.org/abs/1209.2355
  2. Li, L. et al. (2010). Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms. arXiv:1003.5956. https://arxiv.org/abs/1003.5956
  3. Dudík, M., Langford, J., Li, L. (2011). Doubly Robust Policy Evaluation and Learning. arXiv:1103.4601. https://arxiv.org/abs/1103.4601
  4. Ruan, Y. et al. (2023). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817. https://arxiv.org/abs/2309.15817
  5. Jimenez, C. E. et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770. https://arxiv.org/abs/2310.06770
  6. Trivedi, H. et al. (2024). AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
  7. Drouin, A. et al. (2024). WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? arXiv:2403.07718. https://arxiv.org/abs/2403.07718
  8. Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
  9. Swaminathan, A., Joachims, T. (2015). Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. arXiv:1502.02362. https://arxiv.org/abs/1502.02362
  10. Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press. https://doi.org/10.1017/9781108653985
  11. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.