Skip to content

EthenEthenEthen

What Happens When an AI Task Fails Halfway Through?

When an AI task fails halfway through, what happens next depends on one question: did anything change in the world before it failed? A well-designed system handles that question in a fixed order. It stops making new changes. It works out what finished, what did not, and what is uncertain. It checks the uncertain steps by asking the system of record what actually happened. Then it chooses the next move — resume from the last good point, retry only where that is safe, finish just the remaining items, undo a partial change where a real undo exists, or stop and ask you — and it tells you only when a decision genuinely needs you. What it should never do is guess, start over from the beginning, or report the task as either finished or failed when it does not know which.

When an AI task fails halfway through, what happens next depends on one question: did anything change in the world before it failed? A well-designed system handles that question in a fixed order. It stops making new changes. It works out what finished, what did not, and what is uncertain. It checks the uncertain steps by asking the system of record what actually happened. Then it chooses the next move — resume from the last good point, retry only where that is safe, finish just the remaining items, undo a partial change where a real undo exists, or stop and ask you — and it tells you only when a decision genuinely needs you. What it should never do is guess, start over from the beginning, or report the task as either finished or failed when it does not know which.

Key takeaways

  • "Failed halfway" is not one situation. A failure before any action and a failure during a payment need completely different responses.
  • The first move is to stop and take stock. Nothing new should happen until the system knows what already did.
  • Uncertain steps get checked, not repeated. A timeout does not mean the action failed; it means nobody knows yet.
  • Resume, don't restart. Work continues from the last good point, so finished steps are not redone.
  • You should see the truth. Done, not done and unknown — each shown as such, with a decision only when one is needed.

Why is a mid-task failure different from a failed answer?

A failed answer in a chat costs you a few seconds: you ask again. A mid-task failure is different because an AI agent doing real work takes actions with effects — updating records, sending messages, creating files, moving money. By the time it fails, some of those effects may already exist. Asking again from the top would repeat them. Treating the task as failed would leave them in place without anyone knowing.

This is why the honest answer to "what happens?" starts with "where did it fail?".

Table of six failure points with what a good system does and what you should see.
Figure 1. "Failed halfway" is not one situation. The right response depends on whether anything changed in the world.

Failure before any action

If the agent fails while it is still reading, thinking or planning — before it has changed anything — the situation is simple. Nothing in the world is different, so the step can be restarted. You might see a short pause, or nothing at all.

The only thing to get right here is that the work itself is not lost. If the task's progress lives in a browser tab or a single server process, even a harmless failure can lose the plan and force you to start over. Ethen runs long work as durable jobs so that progress survives a crash and another worker can pick the job up; see Inside Ethen's Durable Job Service.

Failure between two actions

If the agent fails after finishing one action and before starting the next, the system should resume from the last completed step. Finished work stays finished; it is not redone. You should see a pause in the task's timeline and then progress continuing.

This only works if the system recorded each completed step durably as it went, rather than reconstructing progress from the agent's memory. A task that "resumes" by asking the model what it thinks it already did is guessing.

Failure during an action

This is the hard case. The agent sent a request — issue this refund, send this message, create this environment — and the confirmation never came back. The request might never have arrived; it might have arrived and failed; or it might have succeeded while the reply was lost. From the agent's side, those three situations look identical.

A well-designed system treats this as an unknown outcome, not as a failure. Retrying immediately risks doing the action twice — a second refund, a duplicate message. Reporting it as failed invites a person to redo it by hand, with the same risk. Instead, the system checks: it asks the payment system, the mailbox or the cloud account whether the action happened. If it did, the step is recorded as done with the evidence. If it did not, retrying is safe. If the truth cannot be established, the system stops and asks a person, with the details needed to check manually.

Ethen's mission system works this way: an action whose outcome is unknown stays marked unknown until an observation resolves it, and the mission cannot complete while it remains. The engineering details are in When an Agent Action's Outcome Is Unknown. Ethen Research Lab's note Unknown Effects in Autonomous AI Systems explains the general rule — reconcile before retry — as a research synthesis.

Retrying safely also depends on the receiving system. Many payment and cloud APIs support idempotency keys: if the same request arrives twice with the same key, the operation happens once (Stripe documentation). The agent's job is to use the same key when it repeats the same intended action, so that a cooperating system can recognize the repeat.

Failure partway through a batch

Many tasks involve many items: migrating 500 records, sending 40 invoices, updating 200 tickets. If the task fails during a batch, some items may have been processed and others not. Restarting the batch duplicates the ones that went through; abandoning it leaves a mess.

The right response is to narrow: find out which items completed, record them, and continue with only the remaining ones. You should see counts — done, remaining, and any that need attention — rather than a vague "partially completed".

Failure because a permission changed

Sometimes a task fails because the agent no longer has the access it needs: an administrator revoked a permission, a token expired, a document's sharing settings changed. Retrying will not help, and working around the missing permission would be the wrong lesson. The right response is to stop that part of the work and ask a person, explaining what access is needed and why. The task waits safely until someone decides.

Failure because a model or provider is unavailable

Sometimes the failure is not in the task's actions but in the AI itself: the model provider is degraded, rate-limited or down. Nothing in the world has changed because of the outage, so this is closer to a failure between actions than during one. A well-designed system can continue with another eligible model if one meets the same requirements and data rules, or pause until the provider recovers.

Two cautions apply. First, a different model may behave differently, so the switch should be recorded and the results reviewed with that in mind. Second, availability never justifies breaking a rule: if the only working alternative would process data somewhere it is not allowed to go, the task should wait. We explain how Ethen thinks about this in Why Ethen Supports More Than One AI Provider.

Failure with no safe way forward

Occasionally there is no safe next move: an outcome cannot be established, the remaining steps depend on something that is broken, or continuing would risk harm. The right response is a safe stop: stop within the task's authority and report exactly what was done, what was not, and what is unknown, so that a person can take over without reconstructing events. A stop with an accurate report is a good outcome. Ethen Research Lab's benchmark design VerifiedWork Recovery proposes scoring agents this way — treating safe, accurately reported stops as acceptable and treating duplicate effects and inaccurate reports as critical failures. It is a benchmark design that has not been run.

Five boxes: Stop new changes, Take stock, Check the uncertain, Choose the next move, Tell you.
Figure 2. The order matters: check before you act again.

Why can't the AI just figure it out?

It is tempting to let the agent reason its way out of a failure — after all, it is a capable model. The problem is that the information it needs is usually not in its context. Whether a payment went through is a fact in the payment system, not something the model can infer. Research has also found that language models struggle to correct their own reasoning without external feedback (Huang et al., 2023). Left alone, an agent will often retry confidently, or write a plausible summary that hides the gap. Recovery has to be built into the system around the agent: durable records of progress, checks against real systems, and hard rules — such as never retrying an action whose outcome is unknown — that the agent cannot talk itself out of.

What should you see as the person who started the task?

You should see the truth, presented simply. For most failures, that means very little: a short pause in the timeline, then progress. For uncertain actions, a status such as "checking what happened" rather than "failed". For failures that need you, a notification with the context to decide — what went wrong, what has already been done, and what your options are. And at the end, a review that shows what was completed, with evidence, and anything that remains open.

We describe how Ethen presents long work in Designing Ethen for Work That Takes Minutes or Hours, and the principles behind recovery in Why Ethen Is Building for Recoverable AI Work.

How can you tell whether a product handles this well?

You can learn a lot about an AI product's reliability by asking how it behaves when a task fails halfway. Five questions are enough.

  1. If the service restarts mid-task, does my task continue, or do I start over?
  2. If an action times out, does the product retry immediately, or check first? Immediate retries on actions like payments or emails are a warning sign.
  3. For batches, does it tell me exactly which items completed? "Partially completed" without counts is not enough.
  4. When it needs me, does it wait safely? It should not proceed on a guess or time out into a default.
  5. After a failure, can I see what was done, what was not and what is unknown? If the only record is the agent's own summary, you are trusting a story.

Products that answer these questions clearly were designed for real work. Products that cannot answer them were designed for demos, where nothing fails halfway.

What happens if I cancel a task halfway?

Cancelling halfway raises the same questions as a failure, because some actions may already have happened and one may be in flight. A good cancel stops new actions immediately, reports which actions were completed, and treats any action that was in flight as uncertain until it is checked. It does not pretend that cancelling undid anything. If some completed actions need to be reversed, that is a separate decision — and for actions that cannot be undone, such as a sent message, the report should say so plainly. We discuss that category in How Ethen Thinks About AI Actions That Can't Be Undone.

A complete example

Illustrative example — hypothetical.

You ask Ethen to send personalized renewal reminders to 120 customers whose contracts expire next month, and to log each one in the CRM. Forty minutes in, the email service starts timing out.

The task stops sending. It takes stock: 73 reminders sent and logged; 46 not yet attempted; one — customer 74 — sent to the email service just as the timeout began, with no confirmation. It checks the email service's sent log for customer 74's message identifier and finds it was sent, so it logs it in the CRM and moves the count to 74. The email service is still timing out, so rather than hammering it, the task pauses and tells you: "Paused: email service unavailable. 74 of 120 sent and logged. Remaining 46 will resume automatically when the service recovers, or you can stop here." An hour later the service recovers and the task finishes the remaining 46. The final review shows 120 sent, 120 logged, no duplicates, and the pause in the timeline.

Frequently asked questions

Will the AI do the same thing twice if it retries? Not in a well-designed system. Before repeating an action whose outcome is uncertain, it checks whether the action already happened, and it uses the same identifier for the same intended action so that cooperating systems can recognize repeats.

Will I lose the work that was already done? No, if the system records progress durably as it goes. The task resumes from its last completed step.

Can I take over a task that failed? You should be able to. A good failure report tells you what was done, what was not, and what is unknown, so you can continue by hand or adjust and resume.

Does "failed" mean nothing happened? Not necessarily. That is exactly why uncertain steps should be shown as "checking what happened" rather than "failed".

References

  1. Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
  2. Stripe. API reference: Idempotent requests. https://docs.stripe.com/api/idempotent_requests
  3. Ethen Blog (2026). Inside Ethen's Durable Job Service. https://upcube.ai/blog/inside-ethens-durable-job-service
  4. Ethen Blog (2026). When an Agent Action's Outcome Is Unknown. https://upcube.ai/blog/when-an-agent-actions-outcome-is-unknown
  5. Ethen Research Lab (2026). Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry. Research note; research synthesis. https://upcube.ai/resources/research/unknown-effects
  6. Ethen Research Lab (2026). VerifiedWork Recovery: Evaluating AI Agents Under Failure and Partial Effects. Benchmark design; not yet run. https://upcube.ai/resources/research/verifiedwork-recovery