Why Ethen Is Building for Recoverable AI Work
Ethen is building for recoverable AI work because failure is a normal part of long-running work, and what happens after a failure decides whether it becomes a brief pause or an incident. Recoverable AI work means that after any interruption — a crash, a timeout, a provider outage, a revoked permission, a person who needs to think — the work is in a known state, and the next move is a deliberate choice: continue from the last good point, reconcile an uncertain outcome, undo a partial effect where a real undo exists, escalate to a person, or stop safely with an accurate account of what happened. Five foundations make that possible: durable state, actions tied to their intent so repeats do not repeat effects, reconciling before retrying, compensation used only where it is meaningful, and clean stopping and hand-over.
Ethen is building for recoverable AI work because failure is a normal part of long-running work, and what happens after a failure decides whether it becomes a brief pause or an incident. Recoverable AI work means that after any interruption — a crash, a timeout, a provider outage, a revoked permission, a person who needs to think — the work is in a known state, and the next move is a deliberate choice: continue from the last good point, reconcile an uncertain outcome, undo a partial effect where a real undo exists, escalate to a person, or stop safely with an accurate account of what happened. Five foundations make that possible: durable state, actions tied to their intent so repeats do not repeat effects, reconciling before retrying, compensation used only where it is meaningful, and clean stopping and hand-over.
Key takeaways
- Failures are normal; incidents are optional. The difference is what the system does next.
- Recoverable means a known state. After any interruption, it should be clear what happened, what did not, and what is unknown.
- Never retry blind. When an outcome is uncertain, find out what happened before doing anything again.
- Undo is not universal. Some effects can be compensated; others — a sent message, a deleted file — cannot.
- Stopping safely is a success. When no safe move remains, a clear stop beats a confident guess.
Why is recovery a product property, not just an engineering detail?
Recovery is a product property because people experience its absence directly. Without it, a crash in the middle of a long task means starting over and wondering what already happened. A timeout during a payment means a duplicate charge or a nervous call to the bank. An agent that hits a wall at step twelve either guesses its way forward or gives up without explaining. Each of these is an engineering failure that shows up as a broken promise to a person.
Recovery matters more for AI agents than for traditional software for two reasons. First, agents choose their own actions at run time, often through tools of varying quality, so failure points are less predictable than in a fixed workflow. Second, language models are not reliable judges of their own mistakes: research has found that models struggle to correct their reasoning without external feedback (Huang et al., 2023). An agent left to its own devices after a failure will often retry with confidence. Recovery therefore has to be designed into the system around the model.
What makes AI work recoverable?
Five foundations make AI work recoverable. Each answers a specific way that failure turns into harm.
1. Durable state
Progress has to live somewhere that survives the failure. If the record of what has been done lives in a browser tab, a single server process or one provider's session, losing that thing loses the work. Durable execution systems in software engineering are built on this idea: a workflow keeps a record of what has happened so it can resume after infrastructure failures (Garcia-Molina and Salem, 1987). Ethen's shared job service is designed so that only one worker acts on a job at a time, a worker that loses its claim cannot overwrite newer progress, and another worker can take over after a crash; see Inside Ethen's Durable Job Service.
Durable state has an important limit that is easy to forget: a checkpoint saves the agent's progress; it does not undo anything the agent did in the outside world. Resuming from a checkpoint taken before an action was sent can send it again. That is why the next three foundations exist.
2. Actions tied to their intent
When an action might be repeated — by a retry, a resume or a second worker — the system needs a way to recognize "this is the same action I already sent". The standard technique is an idempotency key: an identifier that lets the receiving system perform an operation once even if the request arrives twice. Ethen Research Lab's research note Unknown Effects in Autonomous AI Systems adds an agent-specific rule: keys should be derived from the intended effect, not generated fresh for each attempt, because a model that rewords or recomputes a request should not get a new key for the same intent. The note is a research synthesis grounded in distributed-systems practice; it reports no Ethen measurements.
3. Reconcile before retry
When an action's outcome is uncertain — a request was sent and no confirmation came back — the first move is to find out what happened by checking the system of record. If the effect is there, record it and continue. If it is confirmed absent, retrying is allowed. If the truth cannot be established, stop and escalate. Ethen's mission system treats such outcomes as unknown and keeps them open until an observation resolves them; a mission cannot complete while one remains. See When an Agent Action's Outcome Is Unknown.
4. Undo where it is real
Some effects can be offset. A duplicate payment can often be refunded; a record can be restored; a deployment can be rolled back. Long-running database transactions have been structured for decades as sequences of steps, each with a compensating step, so that partial work can be unwound (Garcia-Molina & Salem, 1987). But compensation is not a universal undo. A sent email can be followed by a correction, which is a different and visible act. A file deleted without a backup is gone. Recoverable systems know which category each action falls into before they take it — which is also why the most consequential actions sit behind approvals, as discussed in How Ethen Thinks About AI Actions That Can't Be Undone.
5. Stop safely, hand over cleanly
Sometimes no safe move remains: the outcome cannot be established, a permission has been revoked, the task has become ambiguous. The right response is to stop, record exactly what is known, unknown and pending, and hand the decision to a person. A stop with an accurate account is a successful recovery. A benchmark that scores stopping as failure teaches agents to guess.
What are the choices after a failure?
After a failure, an agent or system has a small set of recovery choices. Ethen Research Lab's proposal Recovery Atlas names seven families: retry unchanged; re-read the state and reconcile; narrow the action (for example, to the records that did not complete); switch to another tool or path; compensate for a partial effect; escalate to a person; or stop safely. The proposal also sets out a rule baseline that a smarter recovery policy would have to beat — including that retrying is not admissible after an unknown outcome on an action that is not safe to repeat, that a missing permission calls for escalation rather than retry, and that stopping safely is always admissible. The Recovery Atlas is a research proposal that has not been built or tested.
The key idea is that recovery is a choice under constraints, not a reflex. A retry is sometimes exactly right. It is never right merely because something failed. Hard constraints — no blind retries after unknown effects, no actions outside the agent's authority — come first, and judgment operates within them.
What can't recovery promise?
Recovery cannot promise that every action happens exactly once across every external system. Distributed systems cannot generally guarantee exactly-once execution across independent services; they achieve "effectively once" behavior only when the receiving system cooperates, for example by honoring idempotency keys (Helland, 2007). Some external tools offer no way to check whether an action happened, and some offer no way to undo it. In those cases the honest answer is to stop and ask, which is sometimes costly.
We also do not claim that recovery is fully automatic. Our engineering posts are explicit that live checks against every external provider are later work; until a system can be checked, its uncertain outcomes stay visibly unresolved rather than being guessed. That is a deliberate trade: an honest backlog of open questions is better than fabricated certainty.
How will Ethen know recovery is working?
Ethen will know recovery is working by testing it deliberately rather than waiting for production failures. Ethen's engineering posts describe tests in controlled settings that inject crashes before and after actions, duplicate deliveries, stale workers and timeouts that hide whether an effect happened. Ethen Research Lab's benchmark design VerifiedWork Recovery proposes going further: inject failures at defined points in an action's lifecycle, record the ground truth the agent cannot see, and score outcomes so that duplicate effects and inaccurate reports are critical failures while safe stops are acceptable. It is a benchmark design that has not been run. A related protocol asks whether recovery knowledge learned on some tools transfers to tools never seen before; it has not been run either.
Questions to ask about any agent system's recovery
Whether you are building an agent system or evaluating one, six questions reveal how recoverable it is.
- Where does progress live? If the answer is "in the conversation" or "in the worker's memory", a crash loses work.
- What happens to an action that was in flight when the failure happened? The right answer involves checking whether it took effect. The wrong answer is "it gets retried".
- How does the system recognize a repeated action? Look for identifiers tied to the intended effect, honored by the systems receiving the action.
- Which actions can be undone, and how does the system know? That knowledge should be declared with the action, not improvised by a model after something goes wrong.
- What does a stop look like? A good stop reports what was done, what was not and what is unknown, and leaves the work resumable by a person.
- How has recovery been tested? Ask for evidence from deliberately injected failures — crashes before and after actions, duplicate deliveries, timeouts — and for the scope of that testing.
Systems that answer these well tend to fail quietly and recover visibly. Systems that answer them poorly tend to fail loudly, or worse, fail silently in ways that are discovered weeks later.
How does this show up for people using Ethen?
For people, recoverable work should feel unremarkable. A long job that hit a server restart shows a short pause in its timeline. A timed-out action shows "checking what happened" rather than "failed". A job that could not safely continue says why, what was done, and what needs a decision. Nothing is silently repeated, and nothing is silently skipped. We describe the user-facing side in Designing Ethen for Work That Takes Minutes or Hours and walk through a concrete failure in What Happens When an AI Task Fails Halfway Through?.
An example
Illustrative example — hypothetical.
An agent is migrating 500 customer records from an old system to a new one in batches of 50. During batch seven, the connection drops after the new system has accepted some records but before it confirmed the batch. A non-recoverable design retries the whole batch, creating duplicates for the records that already arrived. A recoverable design does three things. It marks batch seven's outcome as unknown. It reconciles by asking the new system which of the fifty record identifiers already exist; thirty-one do. It narrows the action, sending only the nineteen missing records with the same intent-based identifiers. Later, a record fails validation in the new system because of a malformed date. Retrying will not help, and guessing the date would corrupt data, so the agent stops that record, continues with the rest, and reports one record needing human attention, with the source value and the error. The migration finishes with 499 records moved, one flagged, and no duplicates.
Tradeoffs
Recoverability costs engineering effort: durable state, intent-based identifiers, reconciliation paths and clear stopping all take work that a demo can skip. Reconciling before retrying is slower than retrying. Safe stops sometimes leave work for people that a reckless retry would have finished correctly by luck. We accept those costs because the failures they prevent — duplicate payments, corrupted records, silent gaps — are the ones that destroy trust in delegated work.
Frequently asked questions
Can an AI agent undo its actions? Sometimes. Actions with a real compensation, such as refunds or rollbacks, can be offset. Others, such as sent messages or permanent deletions, cannot be undone, which is why they deserve approval before they happen.
Should an AI agent retry when something fails? Only when the failure is transient and the action is safe to repeat, or after confirming that a previous attempt did not take effect. A blind retry after an uncertain outcome risks doing the action twice.
What is a safe stop? Stopping within the agent's authority and reporting accurately what was done, what was not, and what is unknown, so a person can decide what happens next.
Does a checkpoint undo an agent's actions? No. A checkpoint saves the agent's progress. External effects that already happened remain and must be checked.
Related reading
- What We're Improving About Reliability Across Ethen
- AI Agents and Autonomous Work
- Recovery Atlas — Ethen Research Lab research proposal.
References
- Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
- Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742
- Helland, P. (2007). Life beyond Distributed Transactions: an Apostate's Opinion. CIDR 2007. https://www.cidrdb.org/cidr2007/papers/cidr07p15.pdf
- Ethen Blog (2026). Inside Ethen's Durable Job Service. https://upcube.ai/blog/inside-ethens-durable-job-service
- Ethen Blog (2026). When an Agent Action's Outcome Is Unknown. https://upcube.ai/blog/when-an-agent-actions-outcome-is-unknown
- Ethen Research Lab (2026). Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop. Research proposal; untested. https://upcube.ai/resources/research/recovery-atlas
- Ethen Research Lab (2026). Unknown Effects in Autonomous AI Systems. Research note; research synthesis. https://upcube.ai/resources/research/unknown-effects
- Ethen Research Lab (2026). VerifiedWork Recovery. Benchmark design; not yet run. https://upcube.ai/resources/research/verifiedwork-recovery