Research Protocol · 2026-10-03 · Context / Skills / Transfer
Testing Whether Recovery Knowledge Transfers Across Tools
An agent that has learned to recover from failures in one system is only useful if that knowledge carries to systems it has never seen. This protocol tests whether it does, and treats knowing when to stop as part of the answer.
Abstract
Agents fail in the middle of work: a call times out after dispatch, a permission is revoked, a schema changes, a batch write half-succeeds. Choosing what to do next is a distinct skill from doing the task, and it can be studied as such. The Recovery Atlas proposes to build verified knowledge about recovery by branching each failure state and verifying every candidate response. The open question is recovery transfer: does knowledge learned from failures in some tools, versions and environments improve recovery on tools, versions, failure mechanisms and environments that were held out? This protocol specifies the experiment. Recovery knowledge is drawn from source failure families and represented as rules, retrieved examples or a learned policy. It is evaluated on held-out targets in which the correct response is sometimes to retry or repair and sometimes to stop, escalate or reconcile before doing anything. Six arms are compared, from generic retry to a frontier agent recovering unaided. The primary outcome is verified recovery without duplicate or unauthorized effects. A recovery policy that acts confidently on an unfamiliar failure where it should have stopped is scored as a critical failure, not a near miss. No recovery-transfer evidence exists yet; a working title that implied results was replaced for that reason.
Why transfer is the right question
Recovery behavior learned on one system is cheap to obtain and easy to overfit. An agent that has seen many timeouts from one payment API may learn that retrying after two seconds works there, a rule that is harmful on an API that is not idempotent. The value of recovery knowledge therefore lies almost entirely in how it behaves on failures that differ from the ones it was learned from.
Existing evidence is suggestive but not decisive. Failure taxonomies and annotated failure corpora show that agent errors cluster into recurring modes and that root causes can be localized: a modular taxonomy across memory, reflection, planning, action and system operations, with a debugging framework that isolates root causes and provides corrective feedback (Zhu et al.); a multi-agent failure taxonomy with fourteen modes (Cemri et al.); and methods that trace errors to the critical step responsible for a failed trajectory (Qi et al.; Zhang et al.). Agents that store and reuse lessons from past attempts improve on repeated tasks (Shinn et al.; Zhao et al.). At the same time, models often cannot correct their own reasoning without external feedback (Huang et al.). None of these studies tests whether recovery knowledge transfers across tools and versions when the cost of a wrong recovery includes real side effects. That is the gap this protocol targets. The underlying failure descriptions come from Toward a Failure Genome of Software Agents.
Design overview
Figure 1 shows the design.
Figure 1. From source failures to verified outcomes on held-out targets. Recovery knowledge is built only from source failure families, represented as rules, retrieved examples or a learned policy, and evaluated on held-out targets where ground truth, including hidden effects, is visible to the verifier but not to the agent. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Source families. Recovery knowledge is built from failures in a set of source tools and environments, using sandbox fault injection, expert-authored scenarios and rights-cleared reconstructions of real failures, as described for the Recovery Atlas. Each failure is recorded with its failure-genome axes: phase, mechanism, trigger, surface, critical step and consequence.
Representation. The knowledge is captured in one of three forms: a rule table keyed on failure class, a store of verified recovery examples retrieved by similarity, or a policy trained on branch outcomes. Each arm uses only source data.
Held-out targets. Targets differ from sources along one axis at a time, and then along several together, as described below.
Verified outcomes. Every recovery attempt is executed in a restorable environment and scored by verifiers that can see ground truth the agent cannot, including whether a timed-out call actually took effect. This follows the design of VerifiedWork Recovery.
Four kinds of held-out target
Figure 2 shows the held-out axes.
Figure 2. Four held-out axes, from easier to harder. Splits hold out whole tools, versions, mechanisms and environments, never random episodes. Transfer is expected to weaken from top to bottom; the study measures by how much. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
- Unseen API. A tool never seen in source data, from a family whose source tools were seen, for example a second ticketing system after training on one.
- Unseen tool version. A new interface version of a source tool, including changes that alter semantics, such as an endpoint that becomes non-idempotent or a field whose meaning changes.
- Unseen failure mechanism. A failure mechanism absent from all source data, such as a partial batch write when sources contained only whole-call failures. This is the hardest and most informative split.
- Unseen environment. A different surface or runtime, for example moving from an API environment to a browser-based one for the same underlying system.
Splits hold out whole mechanisms, tools and versions, never random episodes. A random split would let a policy succeed by recognizing near-duplicates of failures it has already seen.
Arms
- Generic retry. Retry with backoff up to a limit, then stop. The naive baseline.
- Deterministic recovery rules. The admissibility table proposed for the Recovery Atlas: for each failure class, which recovery families are permitted and in what order, with hard constraints such as never retrying a non-idempotent action after an unknown effect until the system of record has been read. Built by engineers who do not see target data.
- Frontier agent, unaided. A strong model given the failure state, the tool documentation and its general instructions, with no recovery knowledge beyond its own.
- Recovery Atlas, rules distilled. Rules induced from source atlas entries, refining the deterministic table.
- Retrieved recovery examples. The frontier agent of arm 3, plus the most similar verified recovery examples from source data, with their outcomes.
- Learned recovery policy. A small policy trained on source branch outcomes to choose among recovery families. This arm is included only if arms 4 and 5 show that source knowledge helps at all; otherwise there is no case for the cost of training.
All arms act through the same runtime, under the same mandate. None may take an action the mandate forbids, and attempted violations are counted.
When the right answer is to stop
A recovery study that rewards only completion will train and select policies that are too willing to act. In a meaningful share of target episodes, the correct response is not to repair the task but to:
- STOP, recording the task as pending with an account of what is known and unknown, because no admissible action can make progress safely;
- ESCALATE, handing the decision to a person, because the next step requires authority the agent lacks or judgment the mandate reserves; or
- RECONCILE, reading the system of record before any further action, because a previous effect is unknown.
These cases are planted deliberately and in known proportion. The handling of unknown effects, where a call may or may not have taken effect, is developed in Unknown Effects; reconciling before retrying is the standard discipline for non-idempotent operations in distributed systems, where safe retries depend on idempotency (RFC 9110) and multi-step work is undone by compensation rather than rollback (Garcia-Molina & Salem).
Figure 3 shows how outcomes are scored.
Figure 3. Scoring: restraint counts as much as repair. Each target episode has a planted correct response. Completing, or correctly stopping, escalating or reconciling, earns credit. An inadmissible action is a critical failure even when it happens to succeed, because the same choice would duplicate an effect elsewhere. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
A policy that retries where it should have reconciled, and happens to succeed, is still scored as an inadmissible action, because the same choice would cause a duplicate effect on another occasion. A policy that stops where repair was safe and available loses completion credit but incurs no critical failure. The asymmetry is intentional.
An illustrative episode
[ILLUSTRATIVE EXAMPLE — a constructed scenario, not an Ethen result.] Source data contain many failures from one invoicing API, whose create-invoice endpoint accepts an idempotency key, so that retrying after a timeout is safe. The held-out target is a second invoicing system, never seen in source data, whose create endpoint has no idempotency key. In the target episode, the agent's create call times out after dispatch. Ground truth, visible only to the verifier, is that the invoice was created.
Arm 1, generic retry, retries and creates a duplicate invoice: a critical failure. Arm 3, the frontier agent unaided, may read the documentation, notice the missing key and query for the invoice before acting, or may retry; the study measures how often each happens. Arm 5 retrieves source examples in which retrying succeeded, which is exactly the wrong lesson for this target unless the retrieved records carry the idempotency property that made retry safe. Arm 2, deterministic rules, reconciles first because the action is non-idempotent and its effect is unknown, finds the invoice and continues.
The episode illustrates why recovery examples need to carry the conditions under which their recovery was valid, not only the action and the outcome, and why a confident, well-precedented action can be the most dangerous one on an unfamiliar tool. It also shows why the primary outcome rewards the reconciling arm here even though every arm might eventually reach a completed task.
Sample size
Critical failures are rare by design, so the study is sized around them. [ILLUSTRATIVE EXAMPLE — arithmetic only.] To bound an arm's critical-failure rate below 1% at 95% confidence with no observed failures requires roughly 300 target episodes of that type for that arm. With four target types, the confirmatory set therefore needs on the order of a thousand episodes per arm, before clustering widens the bounds. A pilot on one target type estimates episode cost and clustering, and the final sizes are fixed before the confirmatory run.
Outcomes and metrics
Primary. Verified recovery without critical failure: the share of target episodes in which the task reaches a verified acceptable end state, either completed or correctly stopped, escalated or reconciled, with no duplicate effect, unauthorized effect or inadmissible action.
Critical failures. Counts of duplicate effects, unauthorized effects and inadmissible actions, with upper confidence bounds. Because these should be rare, a zero count is reported with its bound, never as zero risk (Hanley & Lippman-Hand).
Abstention quality. Precision and recall of stop, escalate and reconcile decisions against the planted ground truth. A policy that escalates everything has perfect recall and is useless.
Cost. Extra model calls, tool calls, wall time and human minutes consumed by recovery.
Transfer ratio. For each arm, performance on held-out targets divided by performance on held-in source-like targets. A ratio well below one means the arm learned its source systems rather than recovery.
Hypotheses
Pre-registered as [PROPOSED TARGET]s:
- H1. Arm 5 (retrieved examples) exceeds arm 3 (frontier unaided) on the primary outcome on unseen-API and unseen-version targets.
- H2. No arm that uses source knowledge has a higher critical-failure rate than arm 2 (deterministic rules) on any target type.
- H3. On unseen failure mechanisms, the advantage of source knowledge over arm 3 is smaller than on unseen APIs and may be absent.
H3 is stated as an expectation of limited transfer. If source knowledge helps even on unseen mechanisms, that would be the most interesting result of the study.
Competing explanations
- Documentation leakage. A frontier agent may recover well on an unseen API because the API's public documentation was in its training data. Targets include private, synthetic APIs with fresh documentation.
- Surface similarity. Retrieved examples may help because target tools resemble source tools superficially. Similarity between source and target is measured and reported, and results are stratified by it.
- Verifier blind spots. A verifier that cannot see a hidden partial effect would under-count critical failures. Environments expose ground-truth state to the verifier, and a sample of episodes is audited by hand.
- Fault-injection artifacts. Injected faults may be more regular than real ones. A subset of targets is reconstructed from rights-cleared real failures, and results are compared between injected and reconstructed targets, an approach related to disciplined fault injection in production engineering (Basiri et al.).
Analysis
Episodes are clustered by task template and tool. Arm comparisons are paired at the episode level, since every arm faces the same failure states. Intervals come from a cluster bootstrap. Comparisons across target types and arms are controlled for false discovery (Benjamini & Hochberg), with H1 and H2 confirmatory. Critical failures are analyzed as counts with exact bounds. Results are reported per target type; an average across target types would hide exactly the variation the study exists to measure.
Safety constraints
All episodes run in restorable sandboxes. Real failures inform targets only through reconstruction, never through replay against live systems, and hazardous effects are never repeated to generate examples. Learned policies from arm 6 are not deployed on the strength of this study alone; at most they become candidates for the gated process described for the Recovery Atlas, with rules as the fallback. The general measurement discipline for transfer claims is set out in the Capability Transfer Ledger.
Limitations
The protocol has not been run, and its hypotheses and thresholds are proposed. Sandbox environments cannot reproduce every property of real systems, especially timing, partial failures across services and human responses. Planted stop, escalate and reconcile cases define correct behavior by the study designers' judgment, which may differ from an organization's own policy. The arms depend on particular models, and a stronger frontier model may close the gap that recovery knowledge fills. The study covers software tools and APIs, not physical systems.
Conclusion
Recovery knowledge is only valuable if it travels. This protocol tests whether it does, on tools, versions, mechanisms and environments it has never seen, and scores restraint as carefully as repair. A result showing that retrieved or learned recovery beats both rules and an unaided frontier agent without raising critical failures would justify the Recovery Atlas. A result showing that it does not would be just as useful, because it would tell builders to invest in rules and reconciliation instead.
FAQ
Why include stopping as a correct answer? Because in many failures no safe action can make progress. Rewarding only completion selects for policies that act when they should not.
Why is retrying-and-succeeding sometimes scored as a failure? If the action was inadmissible, such as retrying a non-idempotent call after an unknown effect, success that time was luck. The same choice will duplicate an effect elsewhere.
What is the hardest split? Unseen failure mechanisms. Recovery knowledge is most likely to help with new tools that fail in familiar ways, and least likely to help with failures of a new kind.
Related research
- Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop — Recovery Atlas.
- VerifiedWork Recovery: Evaluating AI Agents Under Failure and Partial Effects — recovery benchmark.
- The Capability Transfer Ledger: Measuring Whether AI Skills Survive Model Upgrades — transfer methodology.
- Toward a Failure Genome of Software Agents — failure taxonomy.
- Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry — unknown effects constraint.
References
- Zhu, K. et al. (2025). Where LLM Agents Fail and How They can Learn From Failures. arXiv:2509.25370. https://arxiv.org/abs/2509.25370
- Cemri, M. et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657. https://arxiv.org/abs/2503.13657
- Qi, Y. et al. (2026). TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories. arXiv:2608.06346. https://arxiv.org/abs/2608.06346
- Zhang, S. et al. (2025). Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. arXiv:2505.00212. https://arxiv.org/abs/2505.00212
- Shinn, N. et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. https://arxiv.org/abs/2303.11366
- Zhao, A. et al. (2023). ExpeL: LLM Agents Are Experiential Learners. arXiv:2308.10144. https://arxiv.org/abs/2308.10144
- Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
- Fielding, R., Nottingham, M., Reschke, J. (2022). RFC 9110: HTTP Semantics (idempotent methods, §9.2.2). https://www.rfc-editor.org/rfc/rfc9110
- Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742
- Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
- Basiri, A. et al. (2016). Chaos Engineering. IEEE Software 33(3):35–41. https://doi.org/10.1109/MS.2016.60
- Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Skill IR: Toward Model-Independent Agent Capabilities
A research proposal for Skill IR: AI agent skills as portable, testable capability contracts with permissions, verifiers and compatibility records.
- The Capability Transfer Ledger: Measuring Whether AI Skills Survive Model Upgrades
A methods paper on measuring capability transfer: whether AI agent skills keep working across model, tool and task changes, including negative transfer.
- Evidence-Preserving Context: Compressing Agent Memory Without Losing Obligations
A research proposal for context compaction that keeps obligations, permissions, deadlines and evidence out of lossy summaries while cutting an agent's token cost.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.