What “Done” Should Mean for an AI Agent
For an AI agent, "done" should mean that every requirement of the task has been met and that something other than the agent's own report shows it. Precisely: a task is complete when each required obligation is supported by evidence at the level of checking it needs — a passing test, a reconciled record, a confirmed delivery, an approved review — or has been explicitly waived by the person who owns the task. Three refinements make the definition usable. Keep execution success (the steps ran), task success (the outcome was achieved) and business success (it produced value) apart. Treat success as provisional until it can no longer be reversed. And report partial and unknown outcomes as what they are, instead of rounding them up to done.
For an AI agent, "done" should mean that every requirement of the task has been met and that something other than the agent's own report shows it. Precisely: a task is complete when each required obligation is supported by evidence at the level of checking it needs — a passing test, a reconciled record, a confirmed delivery, an approved review — or has been explicitly waived by the person who owns the task. Three refinements make the definition usable. Keep execution success (the steps ran), task success (the outcome was achieved) and business success (it produced value) apart. Treat success as provisional until it can no longer be reversed. And report partial and unknown outcomes as what they are, instead of rounding them up to done.
Key takeaways
- The agent's claim is not evidence. "I've completed the task" is the start of checking, not the end.
- Done means obligations met. List what must be true for the task to be finished; each needs evidence or an explicit waiver from the owner.
- Three kinds of success. Steps running, the task being achieved and the work paying off are different, and arrive at different times.
- Success can be provisional. A merged change can be reverted; a closed ticket can reopen.
- Partial and unknown are honest answers. They are more useful than a false "done".
Why do AI agents say they're done when they're not?
AI agents say they're done when they're not because, in most systems, completion is something the agent declares rather than something the system determines. The agent runs its steps, sees nothing obviously failing, and writes a summary. The summary is fluent and confident regardless of whether every requirement was met.
This failure has a name in Ethen Research Lab's work: false completion — the agent, or a system acting on its report, declares a task complete while some required obligation is unmet. A patch is written but never deployed. A refund is issued but the ticket stays open. A report is delivered without the comparison the customer asked for. Every individual action may have succeeded; the task did not. The research proposal Commitment Graphs defines false completion this way and proposes an experiment to test whether tracking obligations explicitly reduces it. It is a proposal that has not been tested.
Research on agent failures points the same way. A taxonomy built from annotated traces of multi-agent systems groups failures into three categories, one of which is task verification (Cemri et al., 2025). Benchmarks that grade the final state of an environment, rather than the agent's message, exist because the message is unreliable — some also check for unexpected collateral changes the agent did not mention (Trivedi et al., 2024). And on realistic workplace tasks, a benchmark simulating a small software company found that the strongest agent tested completed only around a third of tasks autonomously (Xu et al., 2024). When completion rates are that far from perfect, an agent's own declaration cannot be the measure.
What should "done" mean, precisely?
"Done" should mean that every required obligation is satisfied — supported by evidence at the right level and not contradicted by later evidence — or waived by the person who owns the task.
That definition has three parts worth unpacking.
Obligations. An obligation is a condition that must hold for the task's goal to be met. "Fix the bug" carries obligations: the bug is reproduced, the change makes the reproduction pass, the existing tests still pass, the change is reviewed, and — if the task includes it — it is deployed. Obligations come from the task's definition, from the type of task, and from the person who asked. They should be written down before the work, not reconstructed afterward.
Evidence at the right level. Not all checks are equally strong. Ethen Research Lab's methods paper Evaluating the Evaluators orders verification by trust: deterministic checks first (a test passes, a balance reconciles, a record exists), then programmatic checks, then expert review against a rubric, then calibrated model judges — with the rule that a model's judgment never overrides a failed deterministic check. A refund obligation is best satisfied by the payment system's record; a "the summary is accurate" obligation may need a reviewer. The paper is a research synthesis; it reports no measurements of Ethen's verifiers.
Waivers belong to the owner. Sometimes an obligation cannot be met: the customer is unreachable, a source is behind a paywall. The agent can propose that it be waived — "recommend closing without confirmation" — but only the person who owns the task can waive it. A task closed on an agent's own waiver is a task closed on the agent's say-so.
Execution, task and business success
Separating three levels of success prevents two opposite mistakes.
Execution success means the steps ran and the calls returned. It is known immediately and says little. An email "sent successfully" to the wrong person is an execution success.
Task success means the task's acceptance criteria were met, as judged by a check independent of the agent, and not reversed within a stated window. This is what "done" should usually mean.
Business success means the work produced the value it was for — the customer stayed, the revenue was recovered, the time was saved. It arrives weeks later, if at all, and depends on many things outside the agent's control.
Counting execution success as done inflates results. Holding an agent to business success blames or credits it for things it did not cause. Ethen Research Lab's note The Outcome Warehouse proposes keeping the three apart in how outcomes are recorded; it is an architecture proposal that has not been built.
Why is success sometimes provisional?
Success is sometimes provisional because some outcomes can be reversed after a task ends. A merged change may be reverted next week. A closed support ticket may reopen. A research report may be accepted and later found to rely on a retracted source. Treating these as final the moment the agent stops overstates what was achieved.
A better model treats an outcome as a state that can change. Success is provisional until a window appropriate to the kind of work has closed. Partial success is its own verdict, because a task that is mostly done carries information a simple pass/fail discards. Unknown is distinct from failure. And when new evidence arrives, the old judgment is superseded by a new one rather than overwritten, so anything built on the old judgment can be found and revisited. Ethen Research Lab develops this view in From AI Traces to Verified Experience, a research synthesis.
What should an agent report when it isn't done?
When an agent is not done, the most useful report says exactly what is and isn't true. Four answers are better than a false "done":
- Partially done: which obligations are met, with evidence, and which are not.
- Done, pending confirmation: the work is finished but a required confirmation — a customer reply, a deployment check — has not arrived yet.
- Unknown: an action's outcome could not be established, for example after a timeout, and here is what was done to find out.
- Stopped: the work could not continue safely or within its authority, and here is why.
Each of these leads to a clear next step for a person. A false "done" leads to the wrong one.
What does "done" look like for different kinds of work?
The definition stays the same across kinds of work; the obligations and the evidence change.
- Code. The problem is reproduced; the change makes the reproduction pass; existing tests still pass; the change is reviewed; it is merged or deployed if the task included that. Evidence: test runs, the diff, the review, the deployment record.
- Research. The question is answered; each claim is supported by a cited passage; contradictions are noted; open questions are listed. Evidence: claim-to-source links and a reviewer's sign-off for judgment calls.
- Operations. The requested change exists in the system of record; nothing else changed; the people who needed to know were told. Evidence: the record itself, an audit of collateral changes, delivery confirmations.
- Creative work. The deliverables match the brief and the approved style; rights for any likeness or voice are recorded; the client approved the final version. Evidence: the assets with their lineage, and the approval.
In every case, the agent's summary is useful for orientation and irrelevant as proof.
Will better models make this unnecessary?
Better models will make fewer mistakes per step, and that is valuable. They will not make a definition of done unnecessary, for two reasons. First, more capable agents are given longer and more consequential tasks, so each undetected gap costs more, not less. Second, many failures are not model errors at all: an email bounces, a payment provider times out, a reviewer is on holiday, a requirement was never written down. No model improvement fixes an obligation nobody listed or a confirmation that never arrived. Ethen Research Lab makes the broader argument in Why Better Foundation Models May Make Evaluation More Valuable, Not Less, a position paper and research synthesis.
How does Ethen apply this?
Ethen applies this principle in its mission system: the component that performs a task cannot mark it successful on its own. A task moves into verification with its evidence, and only a separate check can move it to succeeded. A mission completes only when its required tasks have succeeded, its evidence is current, and no action's outcome remains unknown. The mechanism, and the scope of the tests behind it, is described in Making Mission Completion Depend on Evidence.
For a practical checklist you can apply to any agent product — defined outcomes, intervention points, honest recovery and evidence-gated completion — see What Makes an AI Agent Job Verifiable. The two pieces are complementary: that checklist asks whether a system is built to be checked; this article defines what it should be checking for.
How do you write a good definition of done?
A good definition of done is written before the work starts and can be checked by someone other than the agent. Four habits help.
State the outcome, not the activity. "Research three vendors" is an activity. "A table of three vendors with price, contract term and support level, each cell linked to its source" is an outcome.
List obligations, including the easy-to-forget ones. The deployment, the notification to the customer, the update to the ticket, the cleanup of temporary resources.
Name the check for each obligation. The test that must pass, the record that must exist, the person who must approve.
Decide what happens if an obligation cannot be met. Who can waive it, and what the agent should report instead.
This is where delegation to AI becomes more reliable without the AI getting any smarter: the agent's work is aimed at an outcome everyone agreed on, and its completion can be judged against it. We describe how this shapes delegated work in Ethen in What Ethen Work Is Meant to Become.
An example
Illustrative example — hypothetical.
A support agent is asked to resolve a customer's report that an invoice was charged twice. Obligations: the duplicate charge is identified; a refund for the duplicate is issued and confirmed by the payment system; the ticket is updated with the resolution; the customer is notified. The agent finds the duplicate, issues the refund, updates the ticket and writes "Resolved — refund issued, customer notified." A check against the obligations finds that the payment system confirms the refund, the ticket is updated, but the notification email bounced. The honest state is partially done: customer not reached. The task does not close; a person decides whether to try another contact method or waive the notification. Under the agent's own report, the customer would have been left wondering why they were charged twice.
Limitations
Precise definitions of done are easiest for tasks with clear, checkable outcomes. Open-ended work — strategy, design, writing quality — resists deterministic checks and depends on human judgment, which is slower and imperfect. Listing obligations takes effort, and missing one creates a gap no check will catch. Provisional windows delay final answers. These are reasons to apply the discipline in proportion to the stakes, not reasons to skip it.
Frequently asked questions
Is a passing test suite proof that a coding agent is done? It is evidence that the code satisfies those tests. It does not prove the tests were adequate, that the change was reviewed, or that it was deployed if that was required. Each obligation needs its own evidence.
Who decides when an AI agent's task is complete? The checks attached to the task's obligations, and the person who owns the task for anything that requires judgment or a waiver — not the agent's own report.
What is false completion? An agent, or a system acting on its report, declaring a task complete while a required obligation is unmet.
Should AI agents ever report "unknown"? Yes. When an outcome cannot be established, "unknown" with an explanation is more useful and safer than a guess.
Related reading
- What Makes an AI Agent Job Verifiable — the practical checklist.
- Agent Verification and Evaluation
- Why Ethen Shows What It Knows—and What It Doesn't
References
- Cemri, M. et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657. https://arxiv.org/abs/2503.13657
- Trivedi, H. et al. (2024). AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
- Xu, F. F. et al. (2024). TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks. arXiv:2412.14161. https://arxiv.org/abs/2412.14161
- Ethen Blog (2026). Making Mission Completion Depend on Evidence. https://upcube.ai/blog/making-mission-completion-depend-on-evidence
- Ethen Blog (2026). What Makes an AI Agent Job Verifiable. https://upcube.ai/blog/what-makes-an-ai-agent-job-verifiable
- Ethen Research Lab (2026). Commitment Graphs: Why AI Agents Need to Know What Is Still Unfinished. Research proposal; untested. https://upcube.ai/resources/research/commitment-graphs
- Ethen Research Lab (2026). Evaluating the Evaluators: Reward Integrity for AI Agents. Methods paper; research synthesis. https://upcube.ai/resources/research/reward-integrity
- Ethen Research Lab (2026). The Outcome Warehouse. Research note; architecture proposal. https://upcube.ai/resources/research/outcome-warehouse
- Ethen Research Lab (2026). From AI Traces to Verified Experience. Research note; research synthesis. https://upcube.ai/resources/research/verified-experience