Benchmark Design · 2026-10-03 · Evaluation & Verification
VerifiedWork Recovery: Evaluating AI Agents Under Failure and Partial Effects
Most agent benchmarks measure what happens when things go right. Production is mostly about what happens when they don't: a timeout after a payment, a stale record, a half-written batch.
Abstract
Agent benchmarks usually present well-behaved environments: tools respond, state is fresh, actions either succeed or fail cleanly. Deployed agents face the opposite. Tools time out after their effect has already happened, caches return stale data, permissions change mid-task, batch operations complete partially, and APIs drift. How an agent behaves next decides whether a failure becomes a recovery, a safe pause or a duplicated irreversible action. This paper describes VerifiedWork Recovery, the failure track of the Ethen VerifiedWork benchmark. It is an agent failure benchmark built on deliberate fault injection at defined points in an action's lifecycle, with environments that record the ground truth of every external effect the agent cannot see directly. We define eight failure families and the behaviors that count as correct for each. Episodes are scored by comparing final state, action record and the agent's own report against ground truth. Duplicate or unauthorized effects and misreports are treated as critical regardless of task completion. Splits hold out whole failure mechanisms so that scores reflect generalization. The track is a proposed design and has not been run.
Why failure needs its own benchmark
Failures in agent systems are common and consequential. Benchmarks of realistic workplace and application tasks report completion rates well short of reliable automation (Xu et al.; Trivedi et al.). Analyses of agent failures increasingly focus on where in a trajectory the critical error occurred (Qi et al.; Zhang et al.) and on how agents might learn from failures (Zhu et al.). Agents do not reliably correct themselves without external feedback (Huang et al.), and their intermediate uncertainty signals are weak predictors of eventual failure (Li et al.).
Yet most benchmarks inject failure only incidentally. A task fails because the agent made a mistake, not because the environment did something realistic and hostile. Distributed-systems engineering learned long ago that resilience cannot be tested by waiting for faults. It has to be tested by injecting them deliberately and observing the system's response, the practice known as chaos engineering (Basiri et al.). VerifiedWork Recovery applies the same idea to agents.
The central hazard is the ambiguous effect. If an agent issues a refund and the request times out, the refund may or may not have happened. A retry might issue a second refund. A report of "failed" might be false. The correct response is to reconcile against the system of record before doing anything else, and if that is impossible, to stop and report the effect as unknown. The note Unknown Effects in Autonomous AI Systems explains why timeouts are not permission to retry. This track measures whether agents behave accordingly.
Ground truth the agent cannot see
Recovery can be scored only if the benchmark knows what actually happened. Every environment in this track includes a ground-truth recorder: an instrumented layer that logs the true state of each external system at every step, including effects the agent never received confirmation of. The agent sees what a real agent would see, such as timeouts, errors, stale reads and partial responses. The grader sees the truth.
This design separates three things that are usually conflated: what the agent did, what happened, and what the agent says happened. All three are needed. An agent that recovers the right final state but reports the wrong thing has created a reconciliation problem for whoever reads its report.
Where faults are injected
Faults are injected at defined points in the lifecycle of a consequential action (Figure 1).
Figure 1. Where faults are injected in an action's lifecycle. Faults are injected at specific points in the lifecycle of a consequential action. The environment records whether the effect actually happened, so every agent response can be scored against ground truth the agent cannot see directly. Evidence label: EXPERIMENT DESIGN. Source: Ethen benchmark design (proposed).
The lifecycle follows the action states used throughout Ethen's architecture work: proposed, admitted, dispatched, effect applied, acknowledgement returned, receipt recorded, reconciled. Injection points include a stale read before the proposal, a tool error before dispatch, a timeout after dispatch in which the effect has been applied but the acknowledgement is lost, a crash after the effect but before a receipt is recorded, and partial completion of a batch. Injection is controlled per episode, so the same task can be run cleanly and under each fault.
Failure families
Figure 2 lists the eight families in the initial design and the behaviors that count as correct.
Figure 2. Failure families and what correct behavior looks like. Each family specifies how the fault is injected and which responses count as correct. Several families have more than one acceptable response; stopping safely with an accurate report is always acceptable when recovery is not possible within authority. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen benchmark design (proposed).
Several families admit more than one correct response. For a stale read, re-reading before acting is correct; so is acting on a read that the agent explicitly validates. For partial batch completion, completing the remainder and compensating for the completed portion may both be acceptable, depending on the task's obligations, as in the compensating steps of long-lived transactions (Garcia-Molina & Salem). One response is acceptable for every family: stopping safely within authority and reporting accurately what is known, unknown and pending. A benchmark that penalizes safe stopping as failure teaches agents to guess.
The families map onto the taxonomy discussed in Toward a Failure Genome of Software Agents, and the space of recovery responses follows the Recovery Atlas.
Scoring
Each episode is classified by comparing final state, action record and the agent's report against ground truth (Figure 3).
Figure 3. Scoring a recovery episode. Every episode is classified by comparing final state, action record and the agent's own report against ground truth. Critical outcomes, such as duplicate or unauthorized effects, and misreports dominate the score regardless of task completion. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen benchmark design (proposed).
- Critical failure. Any duplicate effect, such as a second refund or deployment, or any action outside the agent's mandate. Critical failures dominate the score regardless of whether the task otherwise completed.
- Misreport. The agent's account of what happened disagrees with ground truth, for example by reporting failure when the effect occurred, or success when it did not.
- Recovered. The final state meets the task's success criteria, with no critical failure and an accurate report.
- Safe stop. The final state does not meet the criteria, but the agent stopped within authority and accurately reported what was pending or unknown.
- Unrecovered failure. None of the above.
The primary metrics are the rates of each class per failure family. Secondary metrics include recovery cost (extra tokens, tool calls and time attributable to the fault), human interventions requested and whether they were appropriate, and the escalation rate. We report critical failure and misreport rates as headline numbers alongside recovery, because they are what an operator most needs to know.
An example episode
[ILLUSTRATIVE EXAMPLE — a design sketch, not a run.] The task: a customer has been double-charged; issue a refund for the duplicate charge and update the ticket. The agent's mandate permits refunds up to the disputed amount without further approval. The injected fault: the payment tool times out after dispatch, while the ground-truth recorder shows that the refund was applied.
Consider four agents. The first retries the refund. The ground-truth recorder shows two refunds: a critical failure, whatever the ticket says. The second reports "refund failed, please retry manually". No duplicate occurred, but the report contradicts ground truth: a misreport, which would likely lead a human to issue the duplicate. The third queries the payment system's transaction history, finds the applied refund, records it and updates the ticket: recovered. The fourth cannot query transaction history because its mandate excludes that tool, so it stops, marks the refund as unknown — requires reconciliation, and escalates: a safe stop. Only the third and fourth behaviors are acceptable. Most benchmarks would score the second agent no worse than the fourth, and could not see the first agent's duplicate at all.
Episode design
Every faulted episode is paired with a clean episode of the same task from the same initial state. The difference between clean and faulted outcomes measures the impact of the fault on each agent, separating recovery skill from general task skill. An agent that fails the clean task tells us nothing about recovery.
Each task–fault pair is run several times, because agent behavior is stochastic and recovery behavior especially so. Reliability is reported with pass^k, the probability of succeeding on all k trials (Yao et al.). For recovery, we also report the worst outcome across trials, since an operator cares whether a duplicate effect can happen at all, not only whether it usually does not.
The recovery overhead of a fault is the additional cost, time and tool calls an agent spends relative to its clean episode. A recovery that succeeds at ten times the clean cost may be acceptable for a rare fault and unacceptable for a common one, so overhead is reported per family.
Reporting rare critical failures
Critical failures should be rare, and rare events are hard to measure. If an agent produces no duplicate effects in n independent episodes of a family, the upper 95% confidence bound on its duplicate rate is approximately 3/n (Hanley & Lippman-Hand). Zero duplicates in 100 episodes is therefore consistent with a true rate near 3%, which may be unacceptable for payment workflows. VerifiedWork Recovery reports upper bounds, not "zero observed". Families with consequential effects receive enough episodes to make those bounds meaningful, and reports state the number of episodes behind every safety claim.
Baselines
Each release reports four baselines on the same episodes:
- No recovery: the agent stops at the first fault.
- Naive retry: repeat the failed action a bounded number of times. This baseline should score badly on unknown-effect families, and the benchmark is designed so that it does.
- Rule baseline: deterministic rules mapping failure classes to recovery actions, of the kind proposed for the Recovery Atlas.
- Prompted frontier recovery: a capable model given the same observations and asked to recover.
A recovery approach is interesting only if it beats the rule baseline on held-out families without raising critical failures.
Splits and generalization
Recovery knowledge is valuable only if it generalizes. Splits therefore hold out whole failure mechanisms, tool versions and task templates, not random episodes. A system that recovers well only from fault types it has seen during development has memorized, not learned. The dedicated protocol for this question is Testing Whether Recovery Knowledge Transfers Across Tools. Sealed episodes are managed separately from development episodes, and canary strings mark every task.
Environment requirements
Fault injection requires environments that can apply effects, lose acknowledgements, crash workers and partially complete batches on purpose, and record the truth throughout. Containerized code environments and simulated enterprise systems both support this, the latter described in Ethen Synthetic Enterprise. Emulated environments can extend coverage to risky actions cheaply (Ruan et al.), but conclusions drawn from emulation must be confirmed in higher-fidelity environments. The core suite and its reporting standard are defined in Ethen VerifiedWork.
What this track cannot prove
A good score on VerifiedWork Recovery shows that an agent handled the injected failure families well in these environments. It does not show that the agent recovers safely from every failure in production, from failure combinations not injected, or from the behavior of real services that the environments approximate imperfectly. In particular, real systems sometimes fail in ways no one has catalogued, and those are exactly the failures that matter most.
Limitations
The track has not been built or run. Environment fidelity under fault injection is hard: real services fail in idiosyncratic ways that simulations may not reproduce. The eight families are a starting set, not a complete taxonomy. Correct-behavior definitions for some families involve judgment about acceptable trade-offs, which requires expert review. Fault injection also tests one fault at a time in the initial design; compound failures are left for later versions.
Conclusion
The most important question about a production agent is not whether it succeeds when everything works. It is what it does when something breaks halfway. VerifiedWork Recovery is designed to answer that question with deliberate faults, hidden ground truth, and scoring that treats duplicate effects and false reports as the serious failures they are.
FAQ
Why is safe stopping scored as acceptable? Because when recovery is not possible within authority, stopping and reporting accurately is the correct behavior. Penalizing it teaches agents to guess, which is how duplicate effects happen.
What is a misreport? An agent's account of what happened that disagrees with ground truth, such as reporting a payment as failed when it actually went through.
Can the track test compound failures? Not in the initial design, which injects one fault per episode. Compound failures are planned for later versions.
Related research
- Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action — umbrella benchmark.
- Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop — Recovery Atlas.
- Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry — unknown effects.
- Toward a Failure Genome of Software Agents — failure taxonomy.
- Testing Whether Recovery Knowledge Transfers Across Tools — transfer of recovery.
References
- Basiri, A. et al. (2016). Chaos Engineering. IEEE Software 33(3):35–41. https://doi.org/10.1109/MS.2016.60
- Xu, F. F. et al. (2024). TheAgentCompany. arXiv:2412.14161. https://arxiv.org/abs/2412.14161
- Trivedi, H. et al. (2024). AppWorld. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
- Qi, Y. et al. (2026). TrajDebug: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories. arXiv:2608.06346. https://arxiv.org/abs/2608.06346
- Zhang, S. et al. (2025). Which Agent Causes Task Failures and When? arXiv:2505.00212. https://arxiv.org/abs/2505.00212
- Zhu, K. et al. (2025). Where LLM Agents Fail and How They Can Learn From Failures. arXiv:2509.25370. https://arxiv.org/abs/2509.25370
- Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
- Li, Z. et al. (2026). Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents. arXiv:2608.29685. https://arxiv.org/abs/2608.29685
- Ruan, Y. et al. (2023). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817. https://arxiv.org/abs/2309.15817
- Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
- Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Work Receipts: A Verifiable Record for Autonomous AI Work
A technical report proposing the Work Receipt: one signed record of authority, actions, effects, verification, cost and rights for every unit of autonomous AI work.
- Evaluating the Evaluators: Reward Integrity for AI Agents
A methods paper on reward integrity for AI agents: verifier false accepts and rejects, abstention, grader drift, expert disagreement and reward hacking.
- Counterfactual Replay for AI Agents
A research proposal for counterfactual evaluation of AI agents: replaying completed tasks under alternative models, tools, context and recovery strategies.
Explained on the Ethen Blog
- When an Agent Action's Outcome Is Unknown
An agent action is interrupted mid-flight: did it happen? Ethen's mission reconciler refuses to guess — it retains the effect as unknown until evidence resolves it.
- What We're Improving About Reliability Across Ethen
Reliability for AI products means more than staying up. For AI work that acts — editing files, sending messages, generating paid media, running for hours — a reliable system must do four things: be available, make sure each real-world effect happens once rather than twice, make sure tasks actually succeed rather than merely report success, and tell people the truth about what is pending, failed or unknown. Ethen's reliability program works on all four layers. Its main parts are safe retries that never duplicate effects, durable jobs that survive crashes and resume cleanly, deliberate fault testing, verification that is budgeted before work begins, release evidence that states exactly what was checked, careful handling of model and provider changes, and honest status everywhere. This article describes that program at a high level and links to the engineering posts behind each part.
- Why Ethen Is Building for Recoverable AI Work
Ethen is building for recoverable AI work because failure is a normal part of long-running work, and what happens after a failure decides whether it becomes a brief pause or an incident. Recoverable AI work means that after any interruption — a crash, a timeout, a provider outage, a revoked permission, a person who needs to think — the work is in a known state, and the next move is a deliberate choice: continue from the last good point, reconcile an uncertain outcome, undo a partial effect where a real undo exists, escalate to a person, or stop safely with an accurate account of what happened. Five foundations make that possible: durable state, actions tied to their intent so repeats do not repeat effects, reconciling before retrying, compensation used only where it is meaningful, and clean stopping and hand-over.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.