Research Paper · 2026-10-03 · Data & Learning Systems
Toward a Failure Genome of Software Agents
Agent failures are usually described with a single label, or not described at all. We propose describing each failure along independent axes, and testing whether that description predicts how to recover.
Abstract
Software agents fail often, and in varied ways: they misread specifications, lose obligations during long tasks, call tools with wrong assumptions, follow injected instructions, mishandle ambiguous effects, and declare success too early. Research on agent failures has grown quickly, producing taxonomies of multi-agent failure modes, datasets for attributing failures to responsible agents and steps, methods for tracing the lifecycle of errors through long trajectories, and checklists for flaws in the benchmarks that measure agents. This research paper surveys that work and proposes a synthesis we call the failure genome: a multi-axis AI agent failure taxonomy in which each failed run is described by its phase, mechanism, trigger, surface, critical step, consequence, detection and recovery outcome, rather than by one category. We argue that independent axes support three uses a single label cannot: predicting which recovery works, tracking whether fixes reduce recurrence, and comparing failure profiles across models. We describe a pipeline for producing genome records, review evidence that automated attribution is still unreliable, and state research questions with their falsifiers. The working title "The Failure Genome of Software Agents" was changed to "Toward…" because no Ethen failure dataset has been measured. The taxonomy is proposed, not validated.
Why failure description matters
Failures carry more information per example than successes. A routine success confirms that something works; a failure with an identified cause shows exactly where a system breaks. Much of the improvement work in agent systems, including fixing harnesses, adding evaluation cases, choosing recovery policies and deciding which model to use for which task, depends on knowing what kind of failure occurred. Yet in practice failures are often recorded as "task failed", perhaps with an error message, and analyzed ad hoc.
Two problems follow. First, without consistent description, an organization cannot tell whether a fix worked: the same underlying failure reappears under a different surface symptom and is counted as new. Second, without description that separates what went wrong from where it surfaced, lessons do not generalize. A dropped obligation looks different in a coding task and a support task, but it is the same failure and probably needs the same remedy.
The broader argument that failures are a high-value form of experience is made in From AI Traces to Verified Experience and What Makes AI Data Defensible?.
What existing work provides
Failure-mode taxonomies. The Multi-Agent System Failure Taxonomy was built from expert analysis of annotated traces across several multi-agent frameworks, validated with high inter-annotator agreement (κ = 0.88), and identifies 14 failure modes in three categories: system design issues, inter-agent misalignment and task verification (Cemri et al.). AgentErrorTaxonomy classifies failures by the module in which they arise, namely memory, reflection, planning, action and system-level operations, and pairs the taxonomy with an annotated dataset and a debugging method that isolates root causes and provides corrective feedback (Zhu et al., 2025b).
Failure attribution. The Who&When dataset annotates failure logs from 127 multi-agent systems with the responsible agent and the decisive error step. Its evaluation is sobering: the best automated method identified the failure-responsible agent 53.5% of the time but the failure step only 14.2% of the time, with some methods below random [EXTERNAL PRIMARY-SOURCE RESULT] (Zhang et al.).
Critical-error tracing. TrajDebug locates the earliest error responsible for a failure in long trajectories. It traces each error's resolution status and terminal impact, so that errors later resolved are distinguished from those that caused failure, and it introduces a benchmark of 486 manually annotated failed trajectories from tool-use and coding settings (Qi et al.).
Uncertainty and timing. Uncertainty signals early in long trajectories are weak predictors of eventual failure. Verbal confidence becomes discriminative only near completion, apparently because agents frequently switch paths mid-trajectory (Li et al.).
Evaluation flaws. Some apparent agent failures and successes are artifacts of the benchmark. The Agentic Benchmark Checklist documents task-setup and reward-design issues that can distort measured performance by up to 100% in relative terms (Zhu et al., 2025a).
Figure 1 summarizes what each source contributes, by its primary focus.
Figure 1. What recent work contributes to each axis. A reading of recent work by its primary focus, based on each paper's stated contribution. No single source covers every axis; the genome proposal is an attempt to combine them, not a claim that they are incomplete for their own purposes. Evidence label: QUALITATIVE MATRIX. Source: Survey of cited literature (abstract-level reading).
The proposal: independent axes
A single category per failure forces choices that lose information. Is an agent that called a refund API twice after a timeout suffering a "tool use" failure, an "execution" failure or a "verification" failure? All three are partly right. We propose instead that each failure record carry values on several independent axes (Figure 2).
Figure 2. A failure record described along independent axes. Instead of one category per failure, a genome record describes each failure along several axes that can vary independently. The combination, not any single label, is what we hypothesize predicts the right recovery. Evidence label: TAXONOMY. Source: Ethen taxonomy proposal.
- Phase: where in the work the failure arose: specification, planning, context and memory, tool use, effect execution, verification and termination, recovery, or authority.
- Mechanism: the specific way it went wrong: acting on stale state, dropping an obligation, following an injected instruction, mishandling an unknown effect, being accepted by a lenient verifier, misreading a tool's semantics.
- Trigger: whether the environment caused the failure, as with a tool error, an interface change or a timeout, or the agent caused it with no environmental fault.
- Surface: the kind of system involved: code, web interface, API or documents.
- Critical step: the earliest step at which the agent had enough information to proceed correctly but did not.
- Consequence severity: from harmless to an irreversible unauthorized effect.
- Detection: who noticed the failure, whether the agent, a verifier, a human or a customer, and when.
- Recovery outcome: what was attempted next and whether it worked.
The name "genome" is a metaphor for a structured description built from a small set of recurring elements. It does not imply that failures are inherited or fixed.
From failed run to record
Figure 3 shows the proposed pipeline.
Figure 3. From a failed run to a genome record. Each stage can be automated in part, but the evidence on automated attribution argues for expert review of the critical step on a sample, and for reporting automated-label accuracy alongside any statistics derived from it. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen taxonomy proposal.
A failed run with a verified outcome is first scanned for suspicious steps: tool errors, retries, policy denials, user corrections and large plan deviations. Critical-step attribution follows, either by counterfactual replay or by a judge model. Replay restores the state before a candidate step, changes the action, and observes whether the outcome changes; the machinery is described in Counterfactual Replay for AI Agents. Given the Who&When results, automated attribution cannot yet be trusted alone, and a sample must be reviewed by experts with agreement measured. The record is then classified on the genome axes, linked to any recovery attempts and their verified outcomes, and clustered with similar records. Each cluster becomes a candidate fix, a new evaluation case or a routing rule.
An illustrative record
[ILLUSTRATIVE EXAMPLE — not an Ethen record.] An agent handling a billing dispute issues a refund; the payment call times out; the agent retries and the customer is refunded twice. A reviewer notices a week later during reconciliation. As a genome record: phase, effect execution; mechanism, mishandled unknown effect; trigger, environmental (timeout after the effect was applied); surface, API; critical step, the retry decision, since at that point the agent knew the outcome was ambiguous; consequence, irreversible duplicate payment, reversed at cost; detection, human reconciliation, seven days later; recovery outcome, none attempted by the agent. A single-label taxonomy would file this as a "tool error". The genome record shows instead that the tool behaved normally, the agent's recovery policy was wrong, and no verifier caught the duplicate. Those are three different fixes: a reconcile-before-retry rule, a policy update, and a reconciliation check added to the task's verifier, as discussed in Unknown Effects in Autonomous AI Systems.
What the axes enable
Recovery selection. If recovery effectiveness depends on mechanism and trigger more than on surface, then a recovery that works for stale state in a CRM should also work for stale state in a code repository. The genome makes that hypothesis testable, and the Recovery Atlas uses genome axes to index recovery outcomes. The VerifiedWork Recovery track injects failures with known mechanisms and triggers, which provides labeled data for exactly this test. External, structured description of failures matters because models often cannot correct their own reasoning without external feedback (Huang et al.).
Recurrence tracking. A fix should reduce the rate of a particular mechanism, not merely of a particular error message. Recurrence rates by mechanism, week over week, measure whether improvement is real.
Cross-model comparison. Different models may fail differently: one drops obligations, another follows injections, a third over-retries. Genome profiles compare models on how they fail, not just how often, which informs both model choice and model change assurance.
Verifier diagnosis. When a failure is detected late, for example by a customer rather than a verifier, the detection axis points to a verifier blind spot. Those cases feed the gold sets described in Evaluating the Evaluators.
Open taxonomy, private instances
The structure of a failure taxonomy is most useful when shared. Shared axes and definitions let organizations and researchers compare results. Failure instances, however, often contain customer data, proprietary code or security-sensitive details, and are private by default. We therefore propose publishing the axes, definitions and annotation guidelines openly, and keeping instance corpora private, sharing only aggregate statistics where rights permit.
Research questions and falsifiers
- Reliability. Can trained annotators label the axes with acceptable agreement, measured by chance-corrected statistics such as Cohen's κ (Cohen)? Falsifier: agreement on the mechanism axis comparable to chance after guideline revision.
- Automation. Can automated labeling reach expert-level agreement on any axis? Given current attribution accuracy, we expect phase and surface to be easier than critical step. Falsifier: no axis reaches acceptable agreement.
- Predictive value. Does the genome predict which recovery succeeds better than a single-label taxonomy does? Falsifier: recovery success is predicted equally well by surface alone.
- Stability. Are mechanism categories stable across models and over time, even as their frequencies change? Falsifier: new frontier models produce mostly failures that fit no existing mechanism.
- Recurrence. Do fixes targeted by mechanism reduce recurrence more than fixes targeted by symptom? Falsifier: no difference.
Annotation guidelines in practice
Multi-axis annotation is harder than assigning one label, and the axes interact. Three guidelines reduce ambiguity. First, the critical step is identified before mechanism is labeled, because the mechanism describes what went wrong at that step. Second, trigger is assigned by checking the environment's ground-truth record: if the environment injected or experienced a fault, the trigger is environmental even if the agent handled it badly; the handling then appears in the recovery axis. Third, when several mechanisms contributed, the record lists them in order of contribution rather than forcing one. Each guideline needs worked examples and periodic recalibration sessions among annotators, with agreement re-measured afterward.
How large a corpus is needed
A taxonomy is only as useful as the frequencies estimated from it. Rare mechanisms, which are often the most consequential, need many failure records before their rates and recovery outcomes can be estimated with useful precision. That is a reason to supplement organic failures with deliberately injected ones from controlled environments, whose mechanism and trigger are known by construction, and to report every frequency with its count.
Limitations
This paper proposes a taxonomy and a pipeline; neither has been validated, and no Ethen failure dataset exists. The axes are informed by the literature but not derived empirically from a large corpus, and they will need revision once applied. Expert annotation is expensive, and automated attribution is not yet reliable. Failures in real deployments often have several interacting causes that a structured record simplifies.
Conclusion
Agent failures are the richest information an agent system produces, and most of it is lost when a failure is recorded as "failed". A multi-axis description, covering phase, mechanism, trigger, surface, critical step, consequence, detection and recovery, preserves what is needed to choose recoveries, verify fixes and compare models. Whether that description is reliable and predictive is an empirical question, and the research questions above say how it would be answered.
FAQ
How is this different from existing failure taxonomies? It combines their contributions, namely categories, responsible components, critical steps and evaluation flaws, as independent axes on each failure record, and adds trigger, consequence, detection and recovery outcome.
Can failure attribution be automated? Partly. Published results show automated methods struggle to identify the failing step, so expert review of a sample remains necessary.
Will the instances be published? The axes and guidelines should be open; instance corpora generally contain private data and stay private.
Related research
- Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop — recovery from classified failures.
- VerifiedWork Recovery: Evaluating AI Agents Under Failure and Partial Effects — recovery benchmark.
- Evaluating the Evaluators: Reward Integrity for AI Agents — verifier failures as a class.
- From AI Traces to Verified Experience — failures as verified experience.
- What Makes AI Data Defensible? — failures as defensible data.
References
- Cemri, M. et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657. https://arxiv.org/abs/2503.13657
- Zhu, K. et al. (2025b). Where LLM Agents Fail and How They Can Learn From Failures. arXiv:2509.25370. https://arxiv.org/abs/2509.25370
- Zhang, S. et al. (2025). Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. arXiv:2505.00212. https://arxiv.org/abs/2505.00212
- Qi, Y. et al. (2026). TrajDebug: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories. arXiv:2608.06346. https://arxiv.org/abs/2608.06346
- Li, Z. et al. (2026). Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents. arXiv:2608.29685. https://arxiv.org/abs/2608.29685
- Zhu, Y. et al. (2025a). Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825. https://arxiv.org/abs/2507.02825
- Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
- Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1). https://doi.org/10.1177/001316446002000104
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Commitment Graphs: Why AI Agents Need to Know What Is Still Unfinished
A research proposal for commitment graphs: tracking intent, obligations, preconditions, effects and evidence so AI agents stop declaring unfinished work complete.
- Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop
A research proposal on AI agent error recovery: branch each failure in a sandbox into candidate recoveries and learn when to retry, reconcile, escalate or stop.
- Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry
A research note on idempotency for AI agents: unknown effects after timeouts, reconcile-before-retry, idempotency keys, compensation and exactly-once limits.
Explained on the Ethen Blog
- Why Not Everything Ethen Researches Needs to Become a Product
The relationship between research and product development at Ethen comes down to one rule: every research question ends in one of three outcomes — build, publish only, or stop — and only one of those is a product. Research becomes a product when five things line up: evidence from results rather than proposals, a real need from people doing real work, a cost of ownership we can sustain as models and data change, clearance on safety, privacy and rights, and a natural place in an existing product. Much valuable research meets some of those and not others. It may produce a method others can reuse, a benchmark, a safeguard inside Ethen, a design principle, or a negative result that saves everyone time. That is why a research publication from Ethen is never a product announcement, and why "publish only" and "stop" are normal outcomes rather than failures. This article explains the rule and how to read Ethen research with it in mind.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.