Research Proposal · 2026-10-03 · Adaptive Intelligence
Commitment Graphs: Why AI Agents Need to Know What Is Still Unfinished
An agent that records only what it did cannot tell whether it is done. We propose recording what it is obligated to achieve, and what evidence would show each obligation is met.
Abstract
A recurring failure of AI agents on long tasks, and the subject of this proposal, is false completion in AI agents: the agent reports success while some requirement remains unmet. A patch is written but never deployed. A refund is issued but the ticket stays open. A report is delivered without the customer's requested comparison. Execution traces do not prevent this, because they record activity rather than obligations. This proposal introduces the commitment graph, a small typed representation linking a task's intent to its obligations, their preconditions, the actions taken, the effects observed and the evidence that each obligation was satisfied. A task is complete only when every required obligation is supported by evidence at its required verification level or has been explicitly waived by the principal. We define false completion precisely and propose a minimal schema and an obligation lifecycle. We then describe an experiment, with a strong checklist baseline and component ablations, that asks whether structured commitment semantics reduce false completion more than simpler alternatives at equal cost. The commitment graph is an Ethen research proposal. It has not been tested.
The problem: "done" is a claim, not an event
Agents finish tasks by saying they are finished. In a single-turn question answering setting, that is harmless. In multi-step work with external effects, it is the moment where errors become invisible. Consider a support mission whose full intent is to resolve a customer's reported bug. Resolving it requires reproducing the issue, creating a patch, passing tests, obtaining approval, deploying, and receiving the customer's confirmation. An agent that completes the first five steps and then writes "the issue has been resolved" has made a false claim. Every individual action it took succeeded.
Failure analyses of agent systems repeatedly surface this pattern. A taxonomy built from annotated multi-agent traces groups its failure modes into three categories, one of which is task verification (Cemri et al.). A benchmark that simulates a small software company found that the most competitive agent tested completed about 30% of tasks autonomously [EXTERNAL PRIMARY-SOURCE RESULT] (Xu et al., TheAgentCompany). The length of software tasks that AI agents can complete has been rising quickly (Kwa et al.), which lengthens the span over which obligations must be tracked. Benchmarks that grade the final state of an environment, including checks for unexpected collateral changes, rather than the agent's report, make the gap measurable (Yao et al.; Trivedi et al.).
Two further findings sharpen the problem. Constraints that a user states once, such as "do not delete any emails until I confirm", are frequently lost when a long context is compacted. One recent evaluation reports that current compaction methods retain only 17% of such injected session constraints on average [EXTERNAL PRIMARY-SOURCE RESULT] (Wang et al., Lost in Compaction). And an agent's own uncertainty signals are weak predictors of eventual failure partway through a long task (Li et al.). An agent cannot be relied on to notice what it has forgotten or what it has not finished.
Why traces are not enough
An execution trace answers what happened. It records that the agent called the test runner, received a passing result, called the deployment tool, and wrote a message. It does not record that the task required customer acceptance, so the absence of acceptance is not visible in the trace. The trace has no slot for it.
Checklists help. A flat list of required steps lets an agent or a reviewer tick items off. But checklists lack three things that matter in practice:
- Evidence links. A ticked box does not say what showed the step was done, or how reliable that showing was.
- Dependencies. Steps have preconditions. Deployment requires approval; approval requires a passing test; a passing test means little if the reproducer was wrong.
- State beyond done and not done. An obligation can be blocked, waived, violated, or unknown because an effect could not be confirmed.
Proposal: the commitment graph
A commitment graph is a typed directed graph attached to a task. Its node types are:
- Intent: the principal's goal, in the principal's terms.
- Obligation: a condition that must hold for the intent to be met, with a required verification level.
- Precondition: a condition that must hold before an action may proceed.
- Authority: a reference to the mandate that permits the actions.
- Action: an attempted effect, with its lifecycle state.
- Effect: an observed change in an external system, with its source.
- Evidence: an artifact that supports or contradicts an obligation: a test result, an approval record, a receipt from a provider, a customer reply.
The edge types are derives (intent to obligation), requires (obligation or action to precondition), authorized-by, produces (action to effect), supports and contradicts (evidence to obligation), and supersedes.
Figure 1 shows the support mission as a commitment graph.
Figure 1. A support mission as a commitment graph. Six obligations derived from one intent. Five are satisfied by evidence; customer acceptance is still pending, so the mission is not complete, even though every action the agent took succeeded. Example is illustrative. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen research proposal; illustrative example.
The mission state is computed, not asserted:
A task is complete if and only if every required obligation is satisfied, meaning supported by evidence at or above its required verification level and not contradicted by later evidence, or is waived by the principal.
False completion is then a precise event: the agent, or any system acting on its report, declares the task complete when the computed state is not complete. This definition makes the failure measurable in any environment where obligations and evidence can be enumerated.
The lifecycle of an obligation
Figure 2 shows the states an obligation can occupy.
Figure 2. Lifecycle of a single obligation. An obligation becomes satisfied only when evidence at the required verification level supports it. Waivers come from the principal, never from the agent. Violated and unknown are explicit states, not silent omissions. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen research proposal.
Four rules govern the lifecycle. Satisfaction requires evidence. An agent's statement that it did something is an action record, not evidence that an obligation holds. Waivers belong to the principal. An agent may propose that an obligation be waived ("the customer is unreachable; recommend closing without confirmation") but cannot waive it. Unknown is a first-class state. If an action's effect could not be confirmed, for example because a request timed out after dispatch, the dependent obligation is unknown, not failed and not satisfied; see Unknown Effects in Autonomous AI Systems. Later evidence supersedes. A reverted deployment contradicts a previously satisfied obligation, and the mission returns to incomplete.
What a commitment graph enables
If the representation can be populated reliably, several capabilities follow.
False-completion prevention. The runtime can refuse to mark a task complete, or to bill it, while obligations remain open. The Work Receipt can record exactly which obligations were satisfied and on what evidence.
Recovery selection. After a failure, the outstanding obligations and their blocking preconditions narrow what a sensible recovery looks like. The Recovery Atlas uses this information to choose between retrying, reconciling, escalating and stopping.
Context compilation. Obligations, preconditions and authority are exactly the facts that must survive context compression. A commitment graph keeps them structurally outside lossy summaries; see Evidence-Preserving Context.
Diagnosis. When a mission fails, the graph shows which obligation failed, which evidence was missing, and which precondition was violated. It is a natural index for locating the critical step in a failed trajectory.
Design choices we would make first
We propose starting small. The first version should be a typed relational representation stored alongside the task, not a graph database. Obligations should come from three sources: templates attached to task types ("a code change requires passing tests"), the mandate's verifier requirements, and explicit statements by the principal. Whether models can reliably derive obligations from free-form intent is itself a research question, and derived obligations should be shown to the principal for confirmation. The representation should be small enough that populating it costs less than the failures it prevents.
Related approaches
The idea that work should be tracked by obligations rather than by activity is old in systems engineering, even if it is new for agents. Long-lived transactions in databases were decomposed into sagas: sequences of steps, each with a compensating action, so that a partially completed process could be identified and unwound (Garcia-Molina & Salem). Provenance vocabularies such as W3C PROV-O model which activities generated which entities and under whose responsibility, which is close to the action, effect and evidence portion of a commitment graph. Workflow engines track which steps of a defined process are complete.
What differs for agents is that the process is not fixed in advance. An agent chooses its steps at run time, often re-plans, and may take actions that no template anticipated. Reasoning-and-acting agents interleave thought and action without an explicit record of what remains owed (Yao et al., ReAct). A commitment graph is meant to supply the missing invariant layer: whatever plan the agent follows, the obligations and their evidence are fixed by the intent and the principal, not by the agent's narrative of its own progress.
Measuring false completion in practice
Measuring false completion requires knowing the obligations independently of the agent. In a benchmark this is straightforward: task authors write a hidden obligation list and the evidence that satisfies each item, and graders compare the agent's completion claim with the computed state. In real work the obligation list must come from the principal or from reviewed templates. Three practical details matter. First, the false-pending rate, where the agent reports work as unfinished when it is in fact complete, must be measured alongside false completion, since an overly cautious agent is also costly. Second, rates should be reported with confidence intervals and stratified by task family, because a single pooled number can hide a family where the graph helps a great deal and another where it hurts. Third, obligations that are satisfied only by delayed evidence, such as customer acceptance, require an observation window long enough for that evidence to arrive, and runs still inside the window must be reported as censored rather than counted either way.
Research question and experiment
Research question (R03). Does a structured representation of intent, obligations, preconditions, effects and evidence reduce false completion and improve recovery compared with ordinary traces and with a flat checklist, at matched model, tools and context budget?
Figure 3 summarizes the proposed design.
Figure 3. Experiment design for research question R03. Three conditions at matched model, tools and context budget, followed by ablations that remove one component at a time. The checklist arm is the strongest simple baseline: if it matches the graph, the graph is not justified. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed); not yet run.
Conditions. Three main arms: ordinary trace, flat checklist and full commitment graph. All three use the same model, tools, task set and total context budget, so that any gain cannot come from extra tokens. Three ablations each remove one component: evidence links, effect receipts or dependency typing.
Tasks. Matched tasks are drawn from code, support-operations and research-evidence families. They deliberately include delayed acceptance, conflicting requirements, permissions revoked mid-task, partially met objectives and actions with unconfirmable effects. A synthetic enterprise environment is the natural setting; see Ethen Synthetic Enterprise.
Metrics. The primary metric is the false-completion rate. Secondary metrics are missed obligations, correctly reported pending states, recovery success after injected failures, time to diagnosis for a human reviewer, schema and runtime overhead, and inter-annotator agreement on obligation labels. The VerifiedWork Context track measures obligation retention across long contexts.
Analysis. Paired comparisons on shared task instances, with clustering by task family. An independent reviewer who did not design the schema adjudicates false completions.
Decision rule. The graph is justified only if it reduces false completion relative to the checklist by a margin fixed in advance, and the gain is not explained by additional context or compute. The proposal fails if the checklist performs as well, if obligations cannot be populated with acceptable agreement, or if maintaining the representation costs more than the failures it prevents.
Competing explanations
A positive result could have other causes. The graph might help simply because it repeats requirements in the context; the checklist arm controls for this. It might help because evidence links make the agent more cautious generally, which would show up as more pending reports on tasks that were in fact complete; the false-pending rate is measured for this reason. And gains on synthetic tasks might not transfer to real work. That threat can be reduced only by testing on real, rights-cleared tasks once the synthetic results are in.
Limitations
This proposal has not been tested. The obligation templates, state definitions and verification levels are design choices that may not survive contact with real workflows. Populating obligations from natural-language intent is unsolved and may require human confirmation that limits automation. The representation adds overhead to every task. And the false-completion metric depends on enumerating obligations correctly, which is hardest exactly in the open-ended tasks where false completion is most likely.
Conclusion
Agents need to know what remains unfinished, and so do the people who rely on them. A trace cannot provide that knowledge because it has no place for what was required but not done. A commitment graph records obligations and the evidence for each, and computes completion instead of accepting it as a claim. Whether that structure is worth its cost is an empirical question with a clear test: does it beat a good checklist?
FAQ
Is a commitment graph a task planner? No. A planner proposes actions. A commitment graph records what must be true for the task to be finished and what evidence shows it. A planner can use it, but the graph is useful even when a human makes the plan.
Who writes the obligations? Task templates, the principal, and the mandate's verification requirements. Model-derived obligations are proposals until a human confirms them.
Does this need a graph database? Not to start. A small typed relational schema is enough to test the hypothesis.
Related research
- Mandates: Compiling Human Intent Into Bounded Agent Authority — authority vs obligation.
- Work Receipts: A Verifiable Record for Autonomous AI Work — completion evidence lives in receipts.
- Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop — outstanding obligations drive recovery choice.
- Evidence-Preserving Context: Compressing Agent Memory Without Losing Obligations — obligations survive context compression.
- VerifiedWork Context: Measuring What AI Agents Must Remember — benchmark track for retaining obligations.
- Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry — pending effects are unfinished commitments.
References
- Cemri, M. et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657. https://arxiv.org/abs/2503.13657
- Xu, F. F. et al. (2024). TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks. arXiv:2412.14161. https://arxiv.org/abs/2412.14161
- Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Trivedi, H. et al. (2024). AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
- Wang, Z. et al. (2026). Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction. arXiv:2608.11242. https://arxiv.org/abs/2608.11242
- Li, Z. et al. (2026). Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents. arXiv:2608.29685. https://arxiv.org/abs/2608.29685
- Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499
- Garcia-Molina, H., Salem, K. (1987). Sagas. Proc. ACM SIGMOD. https://doi.org/10.1145/38713.38742
- Yao, S. et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629. https://arxiv.org/abs/2210.03629
- W3C (2013). PROV-O: The PROV Ontology. https://www.w3.org/TR/prov-o/
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop
A research proposal on AI agent error recovery: branch each failure in a sandbox into candidate recoveries and learn when to retry, reconcile, escalate or stop.
- Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry
A research note on idempotency for AI agents: unknown effects after timeouts, reconcile-before-retry, idempotency keys, compensation and exactly-once limits.
- Toward a Failure Genome of Software Agents
A research paper proposing a multi-axis AI agent failure taxonomy, the failure genome, built on recent work on failure modes, attribution and critical steps.
Explained on the Ethen Blog
- Making Mission Completion Depend on Evidence
In the Stage-0 mission system, running the work and proving the work are different transitions — and only independent evidence unlocks the second.
- What Makes an AI Agent Job Verifiable
An agent job is verifiable when someone other than the agent can confirm what was supposed to happen, what actually happened, and what remains unknown. This guide turns that idea into a checklist you can apply to any system.
- Why Ethen Research Lab Publishes Its Work in Public
Ethen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.