Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Trust & Accountable AI Work

Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry

Publication type
Research Note
Research program
Trust & Accountable AI Work
Published
Authors
Ethen Research Lab
Reading time
12 min read

When an agent's request times out, the world may already have changed. Retrying can turn one refund into two. The first move after an ambiguous outcome must be to find out what happened.

Cover image for "Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry". Decorative abstract motif; contains no data.

Abstract

AI agents increasingly take actions with real effects: payments, messages, record updates, deployments. Every such action passes through a network and a remote system, and sometimes the agent does not learn what happened. A request times out, a worker crashes after dispatch, an acknowledgement is lost. In these cases the effect may or may not have occurred, and the agent's observations cannot tell which. Distributed-systems engineering has long known how to handle this: idempotency, deduplication, reconciliation and compensation. It has also long known that universal exactly-once execution across arbitrary external systems is not available. Agents add new hazards. Models are inclined to retry helpfully, may alter parameters on retry, and may report an unknown outcome as failure. This research note sets out the problem and the established techniques, and proposes rules for idempotency for AI agents: treat unknown as a first-class outcome state, reconcile before retrying, bind idempotency keys and approvals to exact intent, compensate only where compensation is meaningful, and stop and escalate when the truth cannot be established. These are Ethen architecture proposals grounded in systems practice. No Ethen measurements are reported.

The ambiguity

Consider an agent issuing a refund through a payment API. It sends the request and waits. The request times out. Three worlds are consistent with that observation (Figure 1).

An agent sends a request and observes a timeout. Three branches show possible worlds: request never arrived (effect not applied); request arrived and failed (effect not applied); request arrived and succeeded, acknowledgement lost (effect applied). Under each, the consequence of a blind retry: safe, safe, duplicate effect (highlighted).

Figure 1. Three worlds an agent cannot tell apart. After a request times out, the agent's observation is identical in all three worlds. Only the external system knows which one is real. A retry is safe in the first world, harmless in the second, and a duplicate effect in the third. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis; standard distributed-systems reasoning.

In the first, the request never reached the payment system. In the second, it arrived and failed. In the third, it arrived and succeeded, but the response was lost. From the agent's point of view these are identical. A retry is safe in the first two worlds and creates a duplicate refund in the third.

This is not an exotic edge case. Timeouts, dropped connections, worker crashes and lost acknowledgements are routine in distributed systems, and an agent that performs many consequential actions per day will meet them regularly. What is new with agents is who handles them. Traditionally, careful engineers wrote retry logic for specific integrations. Agents generate actions dynamically, often through tools that wrap APIs of varying quality, and they are guided by models that tend to "try again" when something seems to have gone wrong.

What distributed systems already know

Idempotency. An operation is idempotent if performing it several times has the same effect as performing it once. HTTP defines idempotent methods, such as PUT and DELETE, whose repeated application is intended to have the same effect as a single request, and notes that clients may retry them automatically after a communication failure (RFC 9110). Setting a field to a value is naturally idempotent. Incrementing a balance or sending a message is not.

Idempotency keys. For non-idempotent operations, many APIs accept a client-supplied key: if the same key is presented twice, the server performs the operation once and returns the original result. Payment APIs commonly provide this (Stripe documentation), and engineering guidance describes idempotency tokens as the standard way to make retries safe (Amazon Builders' Library).

Reconciliation. When the outcome is uncertain, ask the system of record. A payment system's transaction history, an email server's sent folder or a deployment system's state can establish what happened.

Compensation. Where an effect cannot be prevented from happening twice, it may be undone or offset by a compensating action. Long-lived transactions can be structured as sagas: sequences of steps, each paired with a compensation, so that partial executions can be unwound (Garcia-Molina & Salem).

The limits of exactly-once. Distributed systems cannot generally guarantee that an operation across independent systems happens exactly once. They can achieve effectively once behavior when the receiving system cooperates, through idempotency or deduplication. Helland's analysis of systems that avoid distributed transactions makes this point: messages may be delivered more than once, and the receiving side must be designed to tolerate duplicates. An agent platform that promises exactly-once execution across arbitrary third-party APIs is promising something the APIs themselves may not support.

What is new with agents

Four features of agents make ambiguous effects more hazardous than in conventional integrations.

Helpful retries. When a tool returns an error or a timeout, the natural next step for a language model is often to try again, perhaps with slightly different wording or parameters. Agents do not reliably correct their reasoning without external feedback (Huang et al.), and without explicit instruction a model has no particular reason to suspect that its previous attempt succeeded.

Parameter drift on retry. A retry generated by a model may not be the same request. The amount may be recomputed, a recipient list reordered, a message reworded. Even if the API supports idempotency keys, a key generated per call rather than per intent does not protect against a semantically identical action issued as a new call. And a retry with altered parameters may no longer match what was approved.

Misreporting. An agent may report a timed-out action as "failed", leading a human to redo it. Benchmarks that grade on the agent's final message cannot detect this. Benchmarks that grade the environment's state, including unexpected collateral changes (Trivedi et al.), and compare it with the agent's report can.

Crashes and resumption. Long-running agents checkpoint and resume. A checkpoint records the agent's state but does not undo external effects. An agent resumed from a checkpoint taken before a dispatched action may dispatch it again.

Proposed rules

We propose five rules, drawn from Ethen's internal architecture work and standard practice.

1. Unknown is an outcome state. Every consequential action moves through explicit states: proposed, admitted, dispatched, and then applied, failed or unknown. A timeout after dispatch sets the state to unknown, never to failed. The Work Receipt records unknown outcomes as such, and the commitment graph treats dependent obligations as unresolved.

2. Reconcile before retry. After an unknown outcome, the agent's first action is to query the system of record, by idempotency key, correlation identifier or a search for the expected effect (Figure 2). If the effect is found, record it as applied and continue. If it is confirmed absent, a retry is permitted. If the truth cannot be determined, stop, record the action as unknown and escalate. Never retry blind.

Decision flow after a timeout or crash: mark the action unknown; query the system of record by idempotency key or correlation ID. Three outcomes: effect found, so record it as applied and continue; effect confirmed absent, so retry with the same key and same approved parameters; cannot determine, so stop, record unknown, escalate to a human, and do not retry. A side note: if a retry would change parameters, a new approval is required.

Figure 2. Reconcile before retry. After an ambiguous outcome the first move is to ask the system of record, not to act again. A retry is allowed only if the effect is confirmed absent, and then only with the same idempotency key and the same approved parameters. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal.

3. Keys bind intent, not calls. Idempotency keys are derived from the task and the exact intended effect, for example a hash of the task identifier, action type and material parameters, not generated fresh for each call. A retry of the same intent reuses the same key, so a cooperating API deduplicates it. A model that rewords a request does not get a new key for the same intent.

4. Approvals bind exact actions. An approval applies to the exact parameters approved and to a hash of the material content. A retry that changes anything material is a new action and needs a new approval. This closes the gap where an approved action is retried in modified form without anyone noticing. The authority model is described in Mandates.

5. Fence stale workers. When work is resumed or reassigned, the previous worker's authority to perform consequential actions is revoked using a fencing token or epoch. A worker that wakes up late cannot dispatch an action that its replacement may already have dispatched.

A worked example

[ILLUSTRATIVE EXAMPLE — a design sketch.] An agent is resolving a duplicate charge. Its mandate allows refunds up to the disputed amount. It derives an idempotency key from the task identifier, the action type "refund" and the material parameters: the charge identifier and the amount. It records the action as proposed, then admitted by policy, then dispatched. The payment API times out.

The action becomes unknown. The agent does not retry. It queries the payment system for operations carrying its idempotency key and finds none. Because the payment system's records are eventually consistent, it waits for the documented consistency window and queries again. The refund appears, so the action is recorded as applied, with the provider's transaction identifier as evidence, and the agent updates the ticket.

Had the second query also found nothing, the agent would have treated the refund as confirmed absent and retried with the same key and the same amount, so that a delayed first request would be deduplicated by the provider. Had the agent's next attempt recomputed the amount to include a fee adjustment, the change would have made it a new action needing new approval, not a retry. Had the payment system offered no way to query past operations, the agent would have stopped, marked the refund unknown and escalated to a person with the transaction details needed to check manually. At no point does the agent report "refund failed".

Compensation is not always available

Compensation is sometimes treated as a universal undo. It is not (Figure 3).

Matrix of seven action types (issue refund or payment, send email or message, update a record field to a value, append a comment or log entry, delete a file or record, deploy a release, create a ticket) against four columns: naturally idempotent, idempotency key commonly available, compensation possible, recommended policy after timeout. Sending an email is not idempotent and cannot be compensated, so the policy is reconcile via sent items or stop; setting a field to a value is naturally idempotent and can be retried.

Figure 3. Action types, idempotency and compensation. What an agent can safely do after an ambiguous outcome depends on the action. Where an action can be neither safely repeated nor undone, the only safe policies are reconciliation or stopping. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; typical behavior varies by provider.

A duplicate payment can often be reversed, with cost and delay. A record field set twice to the same value needs no compensation. An email sent twice cannot be unsent; a correction can be sent, which is a different and visible act. A deleted file without soft deletion cannot be restored. For actions that can be neither safely repeated nor undone, the only safe responses to ambiguity are reconciliation and stopping. Agents need to know which category each action falls into. That knowledge belongs in the action's declaration, for example in a Skill IR contract, not in the model's judgment at the moment of failure.

Recovery policy and evaluation

These rules define a rule baseline for recovery. The Recovery Atlas asks whether learned recovery policies can do better within the same constraints. Learned policies may choose among admissible actions, but they never override the prohibition on blind retries after unknown effects. The VerifiedWork Recovery benchmark track injects exactly these ambiguities, records the ground truth, and scores duplicate effects and misreports as critical failures. Emulated sandboxes can extend such testing cheaply to risky actions that are hard to stage for real (Ruan et al.), with conclusions confirmed in higher-fidelity environments.

Measuring how often this happens

How often agents meet ambiguous outcomes, and how often they handle them correctly, depends on the tools, providers and workloads involved. We propose tracking four quantities in any deployment: the rate of unknown outcomes per thousand consequential actions, by tool; the share resolved by reconciliation versus escalated; the rate of duplicate effects detected after the fact, by reconciling against systems of record; and the rate of misreports, in which an agent's account disagrees with the reconciled outcome. None of these has been measured for Ethen systems.

Limitations

The rules depend on cooperation from external systems. If an API offers no idempotency key, no way to query past operations and no reversal, the agent's only safe option after an ambiguous outcome is to stop, which may be costly. Reconciliation itself can be ambiguous when systems of record are eventually consistent. Deriving keys from intent requires a stable definition of "the same intent", which is harder for free-form actions than for structured ones. And none of this protects against an agent that intends the wrong action in the first place; it protects only against performing a correct action the wrong number of times.

Conclusion

A timeout tells an agent that it does not know what happened. It does not tell the agent to try again. Treating unknown as a real state, reconciling before retrying, binding keys and approvals to exact intent, fencing stale workers and knowing which actions can be undone are old lessons from distributed systems. They need to become defaults for agents, because agents will meet these failures constantly and, left alone, tend to respond by trying again.

FAQ

Is exactly-once execution possible for AI agents? Not in general across arbitrary external systems. Effectively-once behavior is possible when the receiving system supports idempotency or deduplication, combined with reconciliation.

Why not just make every tool idempotent? Many external APIs are not, and some actions, such as sending a message, cannot be made idempotent from the client side. Reconciliation and stopping remain necessary.

What should an agent report after a timeout? That the outcome is unknown, together with what it did to find out, never that the action failed unless that has been confirmed.

References

  1. Fielding, R., Nottingham, M., Reschke, J. (2022). RFC 9110: HTTP Semantics (§9.2.2, Idempotent Methods). https://www.rfc-editor.org/rfc/rfc9110
  2. Stripe. API reference: Idempotent requests. https://docs.stripe.com/api/idempotent_requests
  3. Amazon Builders' Library. Making retries safe with idempotent APIs. https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/
  4. Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742
  5. Helland, P. (2007). Life beyond Distributed Transactions: an Apostate's Opinion. CIDR 2007. https://www.cidrdb.org/cidr2007/papers/cidr07p15.pdf
  6. Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
  7. Trivedi, H. et al. (2024). AppWorld. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
  8. Ruan, Y. et al. (2023). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817. https://arxiv.org/abs/2309.15817

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Infrastructure

    Inside Ethen's Durable Job Service

    Leases decide who may act, envelopes decide who the work belongs to, and reconciliation decides what actually happened.

  • Engineering

    When an Agent Action's Outcome Is Unknown

    An agent action is interrupted mid-flight: did it happen? Ethen's mission reconciler refuses to guess — it retains the effect as unknown until evidence resolves it.

  • Company

    What We're Building Across Ethen: October 2026 Update

    As of October 2026, Ethen's work falls into five areas. We are organizing Ethen into focused apps that share one foundation. We are hardening that foundation for AI work that runs for minutes or hours: durable jobs, completion that depends on evidence, honest handling of unknown outcomes, and approvals tied to specific actions. We are building a model knowledge layer so people can choose models with sources rather than guesses. We are extending Ethen to local, on-device AI through Desktop. And Ethen Research Lab now publishes its research in public, with every paper labeled by evidence status. This update separates what our public posts describe as implemented from what is design direction, and it makes no launch or date commitments.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.