Skip to content

EthenEthenEthen

What We're Improving About Reliability Across Ethen

Reliability for AI products means more than staying up. For AI work that acts — editing files, sending messages, generating paid media, running for hours — a reliable system must do four things: be available, make sure each real-world effect happens once rather than twice, make sure tasks actually succeed rather than merely report success, and tell people the truth about what is pending, failed or unknown. Ethen's reliability program works on all four layers. Its main parts are safe retries that never duplicate effects, durable jobs that survive crashes and resume cleanly, deliberate fault testing, verification that is budgeted before work begins, release evidence that states exactly what was checked, careful handling of model and provider changes, and honest status everywhere. This article describes that program at a high level and links to the engineering posts behind each part.

Reliability for AI products means more than staying up. For AI work that acts — editing files, sending messages, generating paid media, running for hours — a reliable system must do four things: be available, make sure each real-world effect happens once rather than twice, make sure tasks actually succeed rather than merely report success, and tell people the truth about what is pending, failed or unknown. Ethen's reliability program works on all four layers. Its main parts are safe retries that never duplicate effects, durable jobs that survive crashes and resume cleanly, deliberate fault testing, verification that is budgeted before work begins, release evidence that states exactly what was checked, careful handling of model and provider changes, and honest status everywhere. This article describes that program at a high level and links to the engineering posts behind each part.

Key takeaways

  • Uptime is the first layer, not the whole definition. A service can be up and still unreliable if it duplicates effects or declares false success.
  • Retries are where many AI reliability problems start. Ethen treats an uncertain outcome as unknown and checks before retrying.
  • Crashes are normal, so recovery is designed in. Jobs are built to survive worker failures without losing or repeating work.
  • "Done" must be earned. Completion depends on evidence independent of the agent's own report.
  • Status should never be rounded off. Pending and unknown are real states, shown as such.

What does reliability mean for an AI product?

Reliability in an AI product is the degree to which people can depend on what it does and what it says it did. Traditional reliability engineering focuses on availability and correctness of responses. AI products that take actions add two further concerns: whether actions in the world happen exactly as intended — once, not zero times and not twice — and whether the system's account of its own work is accurate.

We find it useful to think of four layers.

Four stacked bands: Available; Effects happen once; Tasks verifiably succeed; Status is honest.
Figure 1. A system can be up and still unreliable if it duplicates effects, declares false success or hides what it does not know.

Available. The service responds when people need it. For AI products this includes routing around degraded model providers — within the rules attached to each request.

Effects happen once. When an AI system pays, sends, deploys, deletes or creates something, that effect should happen exactly as intended. The most common way to violate this is a well-meaning retry after an ambiguous failure.

Tasks verifiably succeed. The work should actually achieve what was asked, and that should be checked by something other than the agent's own summary.

Status is honest. When work is pending, has failed, or has an unknown outcome, people should see exactly that, along with what is being done about it.

The layers depend on each other. High availability achieved by aggressive retries can break the second layer. A system that never retries may protect the second layer but fail the first. The program below tries to improve all four without trading one for another.

Safe retries: never duplicate an effect

Retrying is the oldest reliability technique, and for AI agents it is also one of the most dangerous. When an agent's request times out after it was sent, the action may have happened, may have failed, or may never have arrived. A retry is safe in two of those cases and creates a duplicate — a second refund, a second email, a second deployment — in the third.

Ethen treats that situation as an unknown outcome, not a failure. The first move after an unknown outcome is to find out what happened by checking the system of record; only if the effect is confirmed absent is a retry allowed. In Ethen's mission system, an action whose effect is unknown stays marked that way until an observation resolves it, and a mission cannot complete while such an action is unresolved. The mechanism is described in When an Agent Action's Outcome Is Unknown. Ethen Research Lab's research note Unknown Effects in Autonomous AI Systems sets out the general rules — reconcile before retrying, tie retry keys to the intended effect rather than to each attempt — as a research synthesis grounded in distributed-systems practice.

Durable jobs: survive crashes, resume cleanly

AI work increasingly outlives the process that started it. Workers crash, deployments roll, networks drop. Ethen's shared durable job service is designed for that reality: only one worker may act on a job at a time, a worker that has lost its claim cannot overwrite the work of the worker that replaced it, and jobs whose outcome is uncertain are reconciled from evidence instead of being retried by default. Creating the same job twice with the same identity produces one job, not two. The details are in Inside Ethen's Durable Job Service.

The user-facing side of durability is that work can pause and resume. A long task interrupted by a crash, an outage or a person who needs time to approve something should continue from where it was, without repeating what was already done. We explore that experience in Designing Ethen for Work That Takes Minutes or Hours and the principle in Why Ethen Is Building for Recoverable AI Work.

Verified completion: earn "done"

An agent saying it is finished is the least reliable evidence that it is. In Ethen's mission system, the component that performs a task cannot mark it successful on its own; success requires a separate check with evidence. A mission completes only when its required tasks have passed that check, its evidence is current, and no outcome remains unknown. See Making Mission Completion Depend on Evidence.

A related, easily overlooked failure is running out of budget before the checking happens. An agent that spends its whole allowance doing the work has nothing left to verify it, so the work either completes unchecked or stalls. Ethen's mission system reserves verification capacity before execution spending begins, as described in Reserving a Budget for Verification.

Deliberate fault testing

Reliability cannot be established by waiting for failures to happen in production. Distributed-systems engineering learned long ago to inject faults deliberately and observe how a system responds, the practice known as chaos engineering (Basiri et al., 2016). Ethen's engineering posts describe tests that do exactly this in controlled settings: crashes before and after an action, duplicate deliveries, stale workers waking up late, and timeouts that hide whether an effect happened. Each post states the scope of those tests — including what they did not cover, such as live connections to every external provider.

Ethen Research Lab has gone further and designed a benchmark specifically for this layer. VerifiedWork Recovery proposes injecting failures at defined points in an action's lifecycle and scoring agents on whether they recover, stop safely or create duplicate effects. It is a benchmark design that has not been run.

Table of six failures with why each hurts and Ethen's approach.
Figure 2. Each row corresponds to a published engineering mechanism or practice.

Release evidence that says exactly what it proves

A common reliability failure is not technical but interpretive: a release report is read as a broader claim than it makes. Ethen's release certificates are written as dated scope statements: which checks ran, against which code or deployment, on what date, and with what result — including checks marked partial, not run or not performed. A certificate that passes for one product on one date says nothing about another product or a later date. We explain how to read them in What a Release Certificate Actually Proves at Ethen.

This discipline also governs how we talk about rare failures. Observing no failures in a test is useful evidence, but not proof that failures cannot happen. A classic statistical rule of thumb is that zero events in n independent trials still leaves an upper 95% confidence bound of roughly 3/n on the true rate (Hanley & Lippman-Hand, 1983). Zero duplicate effects in 100 trials is consistent with a true rate near 3%. When we report safety-relevant results, we intend to report the bound, not "zero".

Model and provider changes

AI products have a reliability risk that conventional software does not: the model underneath can change. Providers release new versions, retire old ones, and sometimes change behavior behind a stable name. A change that improves average quality can still break specific workflows.

Ethen manages this at two levels. At request time, the AI Gateway checks eligibility — capability, provider health, budget and policy — before choosing a model, and refuses rather than silently substituting something ineligible; see How Ethen Gateway Chooses an Eligible Model. At decision time, we treat model changes as changes to be tested on real work before relying on them, a practice we describe in Why Ethen Sometimes Won't Use the Newest Model. Alternative paths have to be exercised to be trusted; engineers at large cloud providers have written about how rarely used fallback logic can turn a partial failure into a full one (Gabrielson, Amazon Builders' Library).

Honest status everywhere

The fourth layer is the one users see most directly. Every Ethen surface that shows the state of work should distinguish at least: running, waiting for a person, completed with evidence, failed with a reason, and unknown with what is being done to resolve it. Status that rounds "unknown" to "failed" invites duplicate work; status that rounds it to "succeeded" invites people to build on something that may not exist. The same principle applies to information: Ethen Model Intelligence shows Unknown when a model fact lacks provenance rather than filling the gap. We explain why this matters across the product in Why Ethen Shows What It Knows—and What It Doesn't.

One task through all four layers

Illustrative example — a hypothetical task used to show how the layers interact, not a description of a specific incident.

An operations agent is asked to provision a test environment for a new project, register it in the team's inventory, and post a link in the project channel once it is ready.

Available. The model provider the agent normally uses is degraded. The Gateway re-checks eligibility, and a healthy model that meets the same rules continues the planning. The work does not stall, and the record shows which model was used.

Effects happen once. The request to create the environment times out. The environment may or may not exist. Instead of retrying, the task marks the creation as unknown and checks the cloud account for an environment with the task's identifier. It finds one, already running. The task records the creation as done, with the cloud provider's identifier as evidence, and does not create a second environment — which would have cost money and confused everyone.

Tasks verifiably succeed. Before declaring the work finished, a separate check confirms that the environment responds, that the inventory entry exists and points to it, and that the link posted in the channel matches. The agent's own message that "everything is set up" is not what closes the task; the checks are.

Status is honest. Midway, the worker running the task is restarted during a routine deployment. The task resumes on another worker from its last recorded step, with the environment already marked as created. The person who asked for it sees a timeline: planning, an unknown outcome resolved by checking, a pause during the restart, and a completed task with its evidence — not a confident summary that hides the interesting parts.

Without the second layer, this team would have two environments. Without the third, the agent might have posted a link to the wrong one. Without the fourth, nobody would know either had nearly happened.

What we are improving next

At a high level, the program's next steps follow from the four layers:

  • Consistency across apps. The mechanisms above exist in specific parts of Ethen. Their value comes from applying the same rules everywhere long-running or consequential work happens.
  • Live reconciliation. Checking what actually happened in external systems after an uncertain outcome requires an integration for each kind of system. Our engineering posts are explicit that these live integrations are later work.
  • Clearer status in the interface. The distinctions the system already tracks — pending, unknown, awaiting approval — need to be equally clear to the people looking at the screen.
  • Measured results. Our research protocols describe how recovery and verification should be measured. Running them, and publishing the results with their bounds, is part of the program.

What we are not claiming

This article does not state uptime figures, service-level objectives, incident counts or measured failure rates. Where it describes a mechanism, the linked engineering post states the conditions under which it was tested, and those conditions do not include every live provider or every workload. Reliability is a property we are improving, not one we are declaring finished.

Frequently asked questions

Does Ethen retry failed AI actions automatically? Not blindly. If an action's outcome is unknown — for example, after a timeout — Ethen checks what happened before any retry. Retries are allowed when the effect is confirmed absent or when the action is safe to repeat.

What happens if Ethen crashes during a long task? Ethen's job service is designed so another worker can take over without repeating work that already happened, and so the task can resume from where it was.

How does Ethen know a task actually succeeded? In Ethen's mission system, success requires a check independent of the component that did the work, with evidence attached.

Does Ethen publish uptime or reliability metrics? Not in this article. When we publish measured reliability results, they will state their method, scope and uncertainty.

References

  1. Basiri, A. et al. (2016). Chaos Engineering. IEEE Software 33(3):35–41. https://doi.org/10.1109/MS.2016.60
  2. Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
  3. Gabrielson, J. Avoiding fallback in distributed systems. Amazon Builders' Library. https://builder.aws.com/content/3EuS9Sakq7L3VLQIF3qzfMfke1Y/avoiding-fallback-in-distributed-systems
  4. Ethen Blog (2026). Inside Ethen's Durable Job Service. https://upcube.ai/blog/inside-ethens-durable-job-service
  5. Ethen Blog (2026). When an Agent Action's Outcome Is Unknown. https://upcube.ai/blog/when-an-agent-actions-outcome-is-unknown
  6. Ethen Blog (2026). Making Mission Completion Depend on Evidence. https://upcube.ai/blog/making-mission-completion-depend-on-evidence
  7. Ethen Blog (2026). Reserving a Budget for Verification. https://upcube.ai/blog/reserving-a-budget-for-verification
  8. Ethen Blog (2026). What a Release Certificate Actually Proves at Ethen. https://upcube.ai/blog/what-a-release-certificate-actually-proves-at-ethen
  9. Ethen Research Lab (2026). Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry. Research note; research synthesis. https://upcube.ai/resources/research/unknown-effects
  10. Ethen Research Lab (2026). VerifiedWork Recovery: Evaluating AI Agents Under Failure and Partial Effects. Benchmark design; not yet run. https://upcube.ai/resources/research/verifiedwork-recovery