Skip to content

EthenEthenEthen

Research Proposal · 2026-10-03 · Adaptive Intelligence

Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop

Publication type
Research Proposal
Research program
Adaptive Intelligence
Published
Authors
Ethen Research Lab
Reading time
12 min read

Failures are where agents most need judgment and where training data is scarcest. We propose branching each failure, in a sandbox, into the recoveries that could have followed it, and learning from all of them.

Cover image for "Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop". Decorative abstract motif; contains no data.

Abstract

When an AI agent fails mid-task, it has to choose what to do next: retry, re-read the state of the world, narrow the action, switch tools, compensate for a partial effect, ask a human, or stop. These choices determine whether a failure becomes a recovered success, a harmless pause or a duplicated payment. Yet AI agent error recovery is learned from sparse and biased data. Each logged failure shows only the one recovery that was attempted, and the riskiest failures are too rare to learn from by observation. This proposal describes the Recovery Atlas: a corpus built by restoring the same failure state in a sandbox, executing several candidate recoveries from it, and verifying each outcome. The result is counterfactual recovery data. We describe the data model and a rule baseline that encodes which recoveries are admissible for which failure classes, and we propose an evaluation that holds out whole failure mechanisms and tool versions. We state the safety constraints, including that hazardous effects are never repeated to generate data, and the conditions under which the idea should be abandoned. The Recovery Atlas is an Ethen research proposal and has not been built.

Why recovery deserves its own research program

Agents fail often on long tasks. Benchmarks that grade on realistic workplace or application tasks report completion rates well short of reliability (Xu et al.; Trivedi et al.). What happens after a failure matters as much as the failure itself. Two agents with the same first-attempt success rate can differ sharply in final outcomes if one recovers well and the other retries blindly or gives up.

Self-correction does not come for free. Language models struggle to correct their own reasoning without external feedback (Huang et al.). Iterative self-feedback and verbal reflection help in some settings (Madaan et al.; Shinn et al.), but they depend on signals the agent can observe, and in agent settings the most important signal is often the state of an external system. Recent work on long-horizon agents finds that uncertainty signals early in a trajectory say little about eventual failure. Verbal confidence at the end of a run separates failures much better, which supports deciding whether to restart at the end rather than intervening partway (Li et al.). Studies of where agents fail, and how they can learn from those failures, are beginning to map the space (Zhu et al.).

Recovery is also where safety and reliability meet. The wrong recovery after an ambiguous failure can turn one effect into two: a second refund, a second deployment, a second email. The note Unknown Effects in Autonomous AI Systems develops why a timeout is not permission to retry.

Why logged failures are not enough

Production logs contain failures and the recoveries that followed them. Learning recovery from those logs runs into three problems.

One branch per failure. A log shows what the agent did after a failure, not what would have happened had it done something else. If the agent retried and succeeded, we do not learn whether re-reading the state first would have been cheaper or safer. This is the counterfactual problem familiar from learning with logged bandit feedback (Bottou et al.).

Biased branches. The recovery chosen in production was chosen by the existing policy. If that policy always retries transient errors, the data contain almost no examples of switching tools for them. A learner trained on such data inherits the policy's blind spots.

Rare, dangerous failures. The failures that matter most, such as ambiguous effects on irreversible actions, are the rarest. Waiting to observe enough of them in production means waiting for harm.

Proposal: branch the failure, verify every branch

The Recovery Atlas addresses all three problems with one mechanism (Figure 1). For each failure, the environment is snapshotted at the point just before the agent's choice of recovery. The snapshot is then restored once per candidate recovery, each candidate is executed, and the outcome is verified.

A failure-state snapshot on the left fans out to seven recovery branches: retry, re-read state, narrow the action, switch tool, compensate, ask for approval, stop safely. Each branch leads to a verified-outcome cell recording completion, duplicate effects, cost and human intervention. One branch, re-read state then retry, is highlighted as the best verified outcome in this illustrative example; blind retry is marked as producing a duplicate effect.

Figure 1. One failure state, seven counterfactual recoveries. In a sandbox, the same pre-divergence state is restored and each candidate recovery is executed and verified. A single logged trajectory shows only the branch that happened; the atlas records all of them. Branch outcomes shown are illustrative. Evidence label: ILLUSTRATIVE — NOT MEASURED ETHEN DATA. Source: Ethen research proposal; illustrative outcomes, not measured.

The candidate set covers seven recovery families:

  • Retry the same action unchanged.
  • Re-read state and reconcile: check the system of record to learn whether the previous action took effect.
  • Narrow the action: reduce its scope, for example to one record instead of a batch.
  • Switch tool: use an alternative path to the same effect.
  • Compensate: issue an action that undoes or offsets a partial effect, in the manner of compensating transactions for long-lived work (Garcia-Molina & Salem).
  • Escalate or ask for approval: hand a decision to a human with an explanation.
  • Stop safely: record the task as pending with a clear account of what is known and unknown.

Each branch is scored on verified completion, duplicate or unauthorized effects, cost, latency, human intervention required and residual risk. The outcome is an atlas entry: the failure state's features, its failure class, the branch taken and the verified outcome of every branch.

Where the failures come from

Atlas entries come from three sources, in order of priority:

  1. Sandbox fault injection. Deliberately induced failures in controlled environments: tool timeouts after dispatch, stale caches, revoked permissions mid-task, changed API schemas, failing tests, partial batch writes. The VerifiedWork Recovery benchmark track defines a family of such failures.
  2. Expert-authored scenarios. Rare failures that are hard to generate automatically, written by engineers or domain experts, with the ownership of the scenarios assigned clearly.
  3. Rights-cleared production failures, reconstructed. A real failure may inform a new scenario only if its rights permit, and only by reconstructing the situation in a sandbox. Production failures are never replayed against live systems to generate branches, and hazardous effects are never repeated to create examples.

Branching requires restorable state. In software environments that means containers, database snapshots and recorded or stubbed external services. The general replay machinery, and its limits for effects that cannot be stubbed, is described in Counterfactual Replay for AI Agents.

A rule baseline that the atlas must beat

A learned recovery policy is worth building only if it beats good rules. Figure 2 shows a proposed rule baseline: a prior over which recovery families are admissible for which failure classes.

Matrix of seven failure classes (transient tool error, stale state, unknown effect after timeout, missing permission, wrong tool assumption, failed tests or verifier, partial completion) against six recovery options (retry, re-read and reconcile, switch tool or narrow action, compensate, escalate or ask approval, stop safely). Retry is marked inadmissible for unknown effects and missing permissions; stop safely is admissible for all classes.

Figure 2. Which recoveries are admissible for which failures. A prior over recovery options by failure class, proposed as the rule baseline that a learned atlas must beat. The most important cell is unknown effect × retry: a blind retry is inadmissible when an irreversible effect may already have happened. Evidence label: QUALITATIVE MATRIX. Source: Ethen research proposal (rule baseline); qualitative.

Several cells encode hard constraints rather than preferences. Retrying is inadmissible after an unknown effect on a non-idempotent action until the system of record has been checked. Retrying is inadmissible after a missing permission; the right move is to escalate or stop. Stopping safely is always admissible, though not always optimal. Within these constraints, the rules order the options by expected cost.

The research question is whether a policy built from atlas entries, either by retrieving similar past failures or by training a small model, chooses better than these rules on failures it has not seen. The failure taxonomy supplies the failure classes. The commitment graph supplies what remains unfinished, which strongly shapes the right recovery.

Evaluation design

Research question (R02). Do counterfactual failure and recovery examples produce recovery gains that transfer to unseen failures and tools, beyond strong rules and frontier-model recovery prompting?

Figure 3 shows the pipeline.

Pipeline: failure sources (sandbox fault injection, expert-authored scenarios, rights-cleared production failures reconstructed in sandbox) feed branch execution with verifiers, producing atlas entries (state features, failure class, branch, verified outcome, cost, risk). Entries are split into training families and held-out mechanisms and tool versions. A recovery policy (retrieved or trained) is evaluated against rules and frontier-model recovery on the held-out split.

Figure 3. From branched failures to a tested recovery policy. Atlas entries are built only in sandboxes or lawful reconstruction environments. Splits hold out whole failure mechanisms and tool versions, so a positive result means recovery knowledge generalized rather than memorized. Evidence label: EXPERIMENT DESIGN. Source: Ethen research proposal.

Arms. No recovery; generic retry; the rule baseline in Figure 2; a frontier model prompted to recover, given the same information; and the atlas policy, either retrieved or trained. All arms share tools, budget and environment.

Splits. Generalization is the point, so splits hold out whole failure mechanisms, tool versions and task templates, not random entries. If the policy improves only on mechanisms it has seen, it has memorized rather than learned. The cross-tool question is the subject of a dedicated protocol, Testing Whether Recovery Knowledge Transfers Across Tools.

Negative cases. The test set includes failures where stopping is correct, so that a policy biased toward action is penalized.

Metrics. Verified completion after recovery; rate of duplicate or unauthorized effects; accepted stops and escalations; human intervention; recovery cost; and transfer across tools. Every critical error, such as a duplicate irreversible effect, is inspected individually by a reviewer independent of the policy's authors.

Gate. A meaningful improvement over both the rules and frontier recovery, with no increase in unacceptable effects, replicated on a second held-out tool or version.

Kill conditions. The idea should be abandoned if gains require copying exact tasks, if gains come only from spending more compute, or if the verifiers used to score branches reward harmful retries.

Why counterfactual data might be valuable

Two arguments suggest the atlas could matter beyond one benchmark. First, counterfactual labels are information that ordinary logs cannot contain at any volume. A production system that executes one recovery per failure will never learn the outcome of the recoveries it did not try. Second, recovery knowledge may be partly model-independent. The fact that a particular payment API returns a timeout while the charge completes is a property of the API, not of the model calling it. Knowledge of this kind may survive model upgrades better than prompt-level tricks. That is a hypothesis, and the capability transfer ledger is how we would measure it.

There is a supportive analogy from robotics. Learning world models from autonomous robot self-play, rather than from success-biased human demonstrations, improved failure prediction and policy evaluation by up to 40%, and improved real-world policy success rates by 65% after reinforcement learning in the learned model [EXTERNAL PRIMARY-SOURCE RESULT] (Yin et al., PlayWorld). Software agents are not robots, and the result says nothing directly about agent recovery. It does suggest that failure-rich, self-generated experience can be more informative than curated successes.

Safety constraints

The atlas exists to reduce harm, so its construction must not cause any:

  • Branches execute only in sandboxes or reconstruction environments with stubbed external effects.
  • Effects that cannot be stubbed safely are excluded rather than approximated with live calls.
  • Emulated environments help explore risky actions cheaply (Ruan et al.), but conclusions drawn from emulation must be checked in higher-fidelity environments before they shape production behavior.
  • Learned recovery policies propose; they never override hard constraints such as the inadmissibility of blind retries after unknown effects, or the authority limits in the agent's mandate.

Limitations

The atlas has not been built and its value is unmeasured. Branching is only as faithful as the sandbox. If the sandbox does not reproduce how a real service behaves under failure, the atlas teaches the wrong lesson. Expert-authored scenarios are costly and may reflect their authors' assumptions. The number of branches grows with the number of candidate recoveries, so cost may limit the corpus to the most important failure classes. And the most valuable failures, such as rare interactions between systems, are exactly those hardest to reproduce faithfully.

Conclusion

An agent's judgment after a failure determines whether that failure is an inconvenience or an incident. Learning that judgment from production logs alone is slow, biased and sometimes dangerous. The Recovery Atlas proposes to generate the missing counterfactuals deliberately and safely: one failure, many verified recoveries, evaluated on failures the policy has never seen. If held-out gains appear, recovery becomes a learnable, transferable capability. If they do not, good rules remain the right answer, and the data will tell us so.

FAQ

Is this reinforcement learning? Not necessarily. The atlas is a dataset of counterfactual outcomes. It can support retrieval of similar past failures, supervised learning of a recovery classifier, or offline reinforcement learning. Which works best is part of the research.

Why not just tell the agent to always check state before retrying? That rule is part of the baseline. The question is whether learned recovery beats such rules on new failures, not whether rules are useful.

Does the atlas use customer data? Only rights-cleared failures, and only after reconstruction in a sandbox. Most entries are expected to come from fault injection and expert scenarios.

References

  1. Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
  2. Madaan, A. et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback. arXiv:2303.17651. https://arxiv.org/abs/2303.17651
  3. Shinn, N. et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. https://arxiv.org/abs/2303.11366
  4. Li, Z. et al. (2026). Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents. arXiv:2608.29685. https://arxiv.org/abs/2608.29685
  5. Zhu, K. et al. (2025). Where LLM Agents Fail and How They Can Learn From Failures. arXiv:2509.25370. https://arxiv.org/abs/2509.25370
  6. Bottou, L. et al. (2012). Counterfactual Reasoning and Learning Systems. arXiv:1209.2355. https://arxiv.org/abs/1209.2355
  7. Xu, F. F. et al. (2024). TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks. arXiv:2412.14161. https://arxiv.org/abs/2412.14161
  8. Trivedi, H. et al. (2024). AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
  9. Yin, T. et al. (2026). PlayWorld: Learning Robot World Models from Autonomous Play. arXiv:2603.09030. https://arxiv.org/abs/2603.09030
  10. Ruan, Y. et al. (2023). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817. https://arxiv.org/abs/2309.15817
  11. Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Engineering

    When an Agent Action's Outcome Is Unknown

    An agent action is interrupted mid-flight: did it happen? Ethen's mission reconciler refuses to guess — it retains the effect as unknown until evidence resolves it.

  • Product

    How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading Paths

    The fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.

  • Engineering

    Why Ethen Is Building for Recoverable AI Work

    Ethen is building for recoverable AI work because failure is a normal part of long-running work, and what happens after a failure decides whether it becomes a brief pause or an incident. Recoverable AI work means that after any interruption — a crash, a timeout, a provider outage, a revoked permission, a person who needs to think — the work is in a known state, and the next move is a deliberate choice: continue from the last good point, reconcile an uncertain outcome, undo a partial effect where a real undo exists, escalate to a person, or stop safely with an accurate account of what happened. Five foundations make that possible: durable state, actions tied to their intent so repeats do not repeat effects, reconciling before retrying, compensation used only where it is meaningful, and clean stopping and hand-over.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.