Skip to content

EthenEthenEthen

Research Protocol · 2026-10-03 · Context / Skills / Transfer

Testing Whether Recovery Knowledge Transfers Across Tools

Publication type
Research Protocol
Evidence status
Protocol / Planned Experiment: The method is defined. The experiment has not been run, so no results are reported.research protocol; not yet run; no recovery-transfer evidence exists
Research program
Context / Skills / Transfer
Published
Authors
Ethen Research Lab
Reading time
13 min read

An agent that has learned to recover from failures in one system is only useful if that knowledge carries to systems it has never seen. This protocol tests whether it does, and treats knowing when to stop as part of the answer.

Cover image for "Testing Whether Recovery Knowledge Transfers Across Tools". Decorative abstract motif; contains no data.

Abstract

Agents fail in the middle of work: a call times out after dispatch, a permission is revoked, a schema changes, a batch write half-succeeds. Choosing what to do next is a distinct skill from doing the task, and it can be studied as such. The Recovery Atlas proposes to build verified knowledge about recovery by branching each failure state and verifying every candidate response. The open question is recovery transfer: does knowledge learned from failures in some tools, versions and environments improve recovery on tools, versions, failure mechanisms and environments that were held out? This protocol specifies the experiment. Recovery knowledge is drawn from source failure families and represented as rules, retrieved examples or a learned policy. It is evaluated on held-out targets in which the correct response is sometimes to retry or repair and sometimes to stop, escalate or reconcile before doing anything. Six arms are compared, from generic retry to a frontier agent recovering unaided. The primary outcome is verified recovery without duplicate or unauthorized effects. A recovery policy that acts confidently on an unfamiliar failure where it should have stopped is scored as a critical failure, not a near miss. No recovery-transfer evidence exists yet; a working title that implied results was replaced for that reason.

Why transfer is the right question

Recovery behavior learned on one system is cheap to obtain and easy to overfit. An agent that has seen many timeouts from one payment API may learn that retrying after two seconds works there, a rule that is harmful on an API that is not idempotent. The value of recovery knowledge therefore lies almost entirely in how it behaves on failures that differ from the ones it was learned from.

Existing evidence is suggestive but not decisive. Failure taxonomies and annotated failure corpora show that agent errors cluster into recurring modes and that root causes can be localized: a modular taxonomy across memory, reflection, planning, action and system operations, with a debugging framework that isolates root causes and provides corrective feedback (Zhu et al.); a multi-agent failure taxonomy with fourteen modes (Cemri et al.); and methods that trace errors to the critical step responsible for a failed trajectory (Qi et al.; Zhang et al.). Agents that store and reuse lessons from past attempts improve on repeated tasks (Shinn et al.; Zhao et al.). At the same time, models often cannot correct their own reasoning without external feedback (Huang et al.). None of these studies tests whether recovery knowledge transfers across tools and versions when the cost of a wrong recovery includes real side effects. That is the gap this protocol targets. The underlying failure descriptions come from Toward a Failure Genome of Software Agents.

Design overview

Figure 1 shows the design.

Flow left to right. Source failure families from fault injection, expert scenarios and reconstructed real failures, each tagged with failure-genome axes. Representation as rules, retrieved examples or a learned policy. Held-out target: unseen API, tool version, failure mechanism or environment. Recovery attempt in a restorable sandbox. Verified outcome: completed, correctly stopped, escalated or reconciled, or critical failure. A wall labeled split by mechanism, tool and version separates source from target.

Figure 1. From source failures to verified outcomes on held-out targets. Recovery knowledge is built only from source failure families, represented as rules, retrieved examples or a learned policy, and evaluated on held-out targets where ground truth, including hidden effects, is visible to the verifier but not to the agent. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Source families. Recovery knowledge is built from failures in a set of source tools and environments, using sandbox fault injection, expert-authored scenarios and rights-cleared reconstructions of real failures, as described for the Recovery Atlas. Each failure is recorded with its failure-genome axes: phase, mechanism, trigger, surface, critical step and consequence.

Representation. The knowledge is captured in one of three forms: a rule table keyed on failure class, a store of verified recovery examples retrieved by similarity, or a policy trained on branch outcomes. Each arm uses only source data.

Held-out targets. Targets differ from sources along one axis at a time, and then along several together, as described below.

Verified outcomes. Every recovery attempt is executed in a restorable environment and scored by verifiers that can see ground truth the agent cannot, including whether a timed-out call actually took effect. This follows the design of VerifiedWork Recovery.

Four kinds of held-out target

Figure 2 shows the held-out axes.

Table with four held-out target types: unseen API in a seen family; unseen tool version including semantic changes; unseen failure mechanism; unseen environment or surface. Columns: what changes, an example, and expected difficulty. Unseen failure mechanism is highlighted as hardest and most informative.

Figure 2. Four held-out axes, from easier to harder. Splits hold out whole tools, versions, mechanisms and environments, never random episodes. Transfer is expected to weaken from top to bottom; the study measures by how much. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

  1. Unseen API. A tool never seen in source data, from a family whose source tools were seen, for example a second ticketing system after training on one.
  2. Unseen tool version. A new interface version of a source tool, including changes that alter semantics, such as an endpoint that becomes non-idempotent or a field whose meaning changes.
  3. Unseen failure mechanism. A failure mechanism absent from all source data, such as a partial batch write when sources contained only whole-call failures. This is the hardest and most informative split.
  4. Unseen environment. A different surface or runtime, for example moving from an API environment to a browser-based one for the same underlying system.

Splits hold out whole mechanisms, tools and versions, never random episodes. A random split would let a policy succeed by recognizing near-duplicates of failures it has already seen.

Arms

  1. Generic retry. Retry with backoff up to a limit, then stop. The naive baseline.
  2. Deterministic recovery rules. The admissibility table proposed for the Recovery Atlas: for each failure class, which recovery families are permitted and in what order, with hard constraints such as never retrying a non-idempotent action after an unknown effect until the system of record has been read. Built by engineers who do not see target data.
  3. Frontier agent, unaided. A strong model given the failure state, the tool documentation and its general instructions, with no recovery knowledge beyond its own.
  4. Recovery Atlas, rules distilled. Rules induced from source atlas entries, refining the deterministic table.
  5. Retrieved recovery examples. The frontier agent of arm 3, plus the most similar verified recovery examples from source data, with their outcomes.
  6. Learned recovery policy. A small policy trained on source branch outcomes to choose among recovery families. This arm is included only if arms 4 and 5 show that source knowledge helps at all; otherwise there is no case for the cost of training.

All arms act through the same runtime, under the same mandate. None may take an action the mandate forbids, and attempted violations are counted.

When the right answer is to stop

A recovery study that rewards only completion will train and select policies that are too willing to act. In a meaningful share of target episodes, the correct response is not to repair the task but to:

  • STOP, recording the task as pending with an account of what is known and unknown, because no admissible action can make progress safely;
  • ESCALATE, handing the decision to a person, because the next step requires authority the agent lacks or judgment the mandate reserves; or
  • RECONCILE, reading the system of record before any further action, because a previous effect is unknown.

These cases are planted deliberately and in known proportion. The handling of unknown effects, where a call may or may not have taken effect, is developed in Unknown Effects; reconciling before retrying is the standard discipline for non-idempotent operations in distributed systems, where safe retries depend on idempotency (RFC 9110) and multi-step work is undone by compensation rather than rollback (Garcia-Molina & Salem).

Figure 3 shows how outcomes are scored.

Decision diagram. A failure state leads to the planted correct response: repair or retry when safe; STOP, ESCALATE or RECONCILE otherwise. The agent's action is compared with it. Matching a correct response without duplicate or unauthorized effects counts as verified recovery. Stopping when repair was safe loses completion credit but is not critical. Any inadmissible action, duplicate effect or unauthorized effect is a critical failure even if the task later succeeds.

Figure 3. Scoring: restraint counts as much as repair. Each target episode has a planted correct response. Completing, or correctly stopping, escalating or reconciling, earns credit. An inadmissible action is a critical failure even when it happens to succeed, because the same choice would duplicate an effect elsewhere. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

A policy that retries where it should have reconciled, and happens to succeed, is still scored as an inadmissible action, because the same choice would cause a duplicate effect on another occasion. A policy that stops where repair was safe and available loses completion credit but incurs no critical failure. The asymmetry is intentional.

An illustrative episode

[ILLUSTRATIVE EXAMPLE — a constructed scenario, not an Ethen result.] Source data contain many failures from one invoicing API, whose create-invoice endpoint accepts an idempotency key, so that retrying after a timeout is safe. The held-out target is a second invoicing system, never seen in source data, whose create endpoint has no idempotency key. In the target episode, the agent's create call times out after dispatch. Ground truth, visible only to the verifier, is that the invoice was created.

Arm 1, generic retry, retries and creates a duplicate invoice: a critical failure. Arm 3, the frontier agent unaided, may read the documentation, notice the missing key and query for the invoice before acting, or may retry; the study measures how often each happens. Arm 5 retrieves source examples in which retrying succeeded, which is exactly the wrong lesson for this target unless the retrieved records carry the idempotency property that made retry safe. Arm 2, deterministic rules, reconciles first because the action is non-idempotent and its effect is unknown, finds the invoice and continues.

The episode illustrates why recovery examples need to carry the conditions under which their recovery was valid, not only the action and the outcome, and why a confident, well-precedented action can be the most dangerous one on an unfamiliar tool. It also shows why the primary outcome rewards the reconciling arm here even though every arm might eventually reach a completed task.

Sample size

Critical failures are rare by design, so the study is sized around them. [ILLUSTRATIVE EXAMPLE — arithmetic only.] To bound an arm's critical-failure rate below 1% at 95% confidence with no observed failures requires roughly 300 target episodes of that type for that arm. With four target types, the confirmatory set therefore needs on the order of a thousand episodes per arm, before clustering widens the bounds. A pilot on one target type estimates episode cost and clustering, and the final sizes are fixed before the confirmatory run.

Outcomes and metrics

Primary. Verified recovery without critical failure: the share of target episodes in which the task reaches a verified acceptable end state, either completed or correctly stopped, escalated or reconciled, with no duplicate effect, unauthorized effect or inadmissible action.

Critical failures. Counts of duplicate effects, unauthorized effects and inadmissible actions, with upper confidence bounds. Because these should be rare, a zero count is reported with its bound, never as zero risk (Hanley & Lippman-Hand).

Abstention quality. Precision and recall of stop, escalate and reconcile decisions against the planted ground truth. A policy that escalates everything has perfect recall and is useless.

Cost. Extra model calls, tool calls, wall time and human minutes consumed by recovery.

Transfer ratio. For each arm, performance on held-out targets divided by performance on held-in source-like targets. A ratio well below one means the arm learned its source systems rather than recovery.

Hypotheses

Pre-registered as [PROPOSED TARGET]s:

  • H1. Arm 5 (retrieved examples) exceeds arm 3 (frontier unaided) on the primary outcome on unseen-API and unseen-version targets.
  • H2. No arm that uses source knowledge has a higher critical-failure rate than arm 2 (deterministic rules) on any target type.
  • H3. On unseen failure mechanisms, the advantage of source knowledge over arm 3 is smaller than on unseen APIs and may be absent.

H3 is stated as an expectation of limited transfer. If source knowledge helps even on unseen mechanisms, that would be the most interesting result of the study.

Competing explanations

  • Documentation leakage. A frontier agent may recover well on an unseen API because the API's public documentation was in its training data. Targets include private, synthetic APIs with fresh documentation.
  • Surface similarity. Retrieved examples may help because target tools resemble source tools superficially. Similarity between source and target is measured and reported, and results are stratified by it.
  • Verifier blind spots. A verifier that cannot see a hidden partial effect would under-count critical failures. Environments expose ground-truth state to the verifier, and a sample of episodes is audited by hand.
  • Fault-injection artifacts. Injected faults may be more regular than real ones. A subset of targets is reconstructed from rights-cleared real failures, and results are compared between injected and reconstructed targets, an approach related to disciplined fault injection in production engineering (Basiri et al.).

Analysis

Episodes are clustered by task template and tool. Arm comparisons are paired at the episode level, since every arm faces the same failure states. Intervals come from a cluster bootstrap. Comparisons across target types and arms are controlled for false discovery (Benjamini & Hochberg), with H1 and H2 confirmatory. Critical failures are analyzed as counts with exact bounds. Results are reported per target type; an average across target types would hide exactly the variation the study exists to measure.

Safety constraints

All episodes run in restorable sandboxes. Real failures inform targets only through reconstruction, never through replay against live systems, and hazardous effects are never repeated to generate examples. Learned policies from arm 6 are not deployed on the strength of this study alone; at most they become candidates for the gated process described for the Recovery Atlas, with rules as the fallback. The general measurement discipline for transfer claims is set out in the Capability Transfer Ledger.

Limitations

The protocol has not been run, and its hypotheses and thresholds are proposed. Sandbox environments cannot reproduce every property of real systems, especially timing, partial failures across services and human responses. Planted stop, escalate and reconcile cases define correct behavior by the study designers' judgment, which may differ from an organization's own policy. The arms depend on particular models, and a stronger frontier model may close the gap that recovery knowledge fills. The study covers software tools and APIs, not physical systems.

Conclusion

Recovery knowledge is only valuable if it travels. This protocol tests whether it does, on tools, versions, mechanisms and environments it has never seen, and scores restraint as carefully as repair. A result showing that retrieved or learned recovery beats both rules and an unaided frontier agent without raising critical failures would justify the Recovery Atlas. A result showing that it does not would be just as useful, because it would tell builders to invest in rules and reconciliation instead.

FAQ

Why include stopping as a correct answer? Because in many failures no safe action can make progress. Rewarding only completion selects for policies that act when they should not.

Why is retrying-and-succeeding sometimes scored as a failure? If the action was inadmissible, such as retrying a non-idempotent call after an unknown effect, success that time was luck. The same choice will duplicate an effect elsewhere.

What is the hardest split? Unseen failure mechanisms. Recovery knowledge is most likely to help with new tools that fail in familiar ways, and least likely to help with failures of a new kind.

References

  1. Zhu, K. et al. (2025). Where LLM Agents Fail and How They can Learn From Failures. arXiv:2509.25370. https://arxiv.org/abs/2509.25370
  2. Cemri, M. et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657. https://arxiv.org/abs/2503.13657
  3. Qi, Y. et al. (2026). TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories. arXiv:2608.06346. https://arxiv.org/abs/2608.06346
  4. Zhang, S. et al. (2025). Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. arXiv:2505.00212. https://arxiv.org/abs/2505.00212
  5. Shinn, N. et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. https://arxiv.org/abs/2303.11366
  6. Zhao, A. et al. (2023). ExpeL: LLM Agents Are Experiential Learners. arXiv:2308.10144. https://arxiv.org/abs/2308.10144
  7. Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
  8. Fielding, R., Nottingham, M., Reschke, J. (2022). RFC 9110: HTTP Semantics (idempotent methods, §9.2.2). https://www.rfc-editor.org/rfc/rfc9110
  9. Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742
  10. Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
  11. Basiri, A. et al. (2016). Chaos Engineering. IEEE Software 33(3):35–41. https://doi.org/10.1109/MS.2016.60
  12. Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.