Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Adaptive Intelligence

From AI Traces to Verified Experience

Publication type
Research Note
Research program
Adaptive Intelligence
Published
Authors
Ethen Research Lab
Reading time
12 min read

An agent's telemetry records what it did. Verified experience also records whether it worked, how we know, what should have happened instead, and whether we may learn from it. The difference decides whether data compounds.

Cover image for "From AI Traces to Verified Experience". Decorative abstract motif; contains no data.

Abstract

Teams that deploy AI agents accumulate logs, traces and trajectories quickly, and it is tempting to treat this accumulation as a strategic asset. This note argues that most agent telemetry is not yet verified experience: a record that ties intent and actions to an outcome judged by a verifier with known error, carries the human correction where one exists, and states the purposes for which it may be reused. We distinguish six representations of the same unit of work and identify the questions each can and cannot answer. We propose treating outcome labels as versioned states rather than booleans, and we describe a retention policy that ranks records by information value instead of volume. The note explains why raw telemetry does not automatically become a data moat, and what has to be added before it can support learning. It is an Ethen Research Lab synthesis. It reports no measured Ethen results.

The accumulation fallacy

An agent platform that records everything will, after a year, hold a large archive of model calls, tool invocations and session transcripts. The archive feels valuable because it is large, unique to the platform and expensive to reproduce. The reasoning runs: competitors have the same models, but not our history.

The reasoning fails at a specific point. A history is valuable for improving a system only if it can say which behaviors were good. Most agent telemetry cannot. It records that a tool returned successfully, not that the customer's problem was resolved. It records that the user stopped responding, not whether the silence meant satisfaction or abandonment. It records what the agent did, not what it should have done. And it rarely records whether the content may be used for anything other than serving that customer. Practitioner experience across agent systems points the same way: traces are cheap and labels are expensive. An unlabeled archive is mostly storage cost and legal surface.

This is not an argument against telemetry. Execution records are necessary for debugging, cost accounting, incident response and audit, and standardization efforts such as the OpenTelemetry semantic conventions for generative AI are making them more interoperable. The argument is narrower. Telemetry is the raw material of verified experience, not a substitute for it.

Six representations

Figure 1 orders six representations of one unit of agent work. Each adds a property the previous one lacks.

Staircase of six boxes from bottom left to top right: log, trace, trajectory, outcome, correction, verified experience. Annotation under each lists what it adds: timestamped events; causal span structure; intent and actions in order; a judged result; the human-preferred alternative; verifier provenance, error rates, rights and cost.

Figure 1. Six representations of the same unit of work. Each step adds a property that the previous representation lacks. Only the last carries everything a learning claim needs: what was intended, what happened, whether it worked according to a verifier with known error, what a human corrected, and whether reuse is permitted. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis.

Log. Timestamped events emitted by components: "model request sent", "tool returned 200". Logs answer what happened, in fragments.

Trace. Events organized into spans with parent–child relations, so a reader can reconstruct causal order across a run. A trace answers what happened, in what order, and in which component.

Trajectory. The sequence of observations and actions taken in pursuit of a task, including the task statement. This is the unit most agent-learning research uses. A trajectory answers what the agent tried to do and how.

Outcome record. A judgment about whether the trajectory achieved its acceptance criteria. Outcomes are where most archives fall short: they are missing, implicit, or inferred from weak proxies.

Correction record. The alternative a human preferred: the edited patch, the rewritten reply, the rejected approval with a reason. Corrections are unusually informative because they show both the error and the fix.

Verified experience. An outcome or correction record with four further properties: (1) the identity, version and measured error rates of the verifier that produced the judgment; (2) the authority under which the work ran; (3) the full cost of the work, including verification; (4) the reuse rights that govern it. Verified experience is the unit that the Verified Adaptive Intelligence agenda learns from, and the Work Receipt is the record format we propose for carrying it.

Questions a learning claim depends on

A claim that a system learned something from deployment rests on six questions. Figure 2 shows which representations can answer them.

Matrix with representations as rows (log, trace, trajectory, outcome record, correction record, verified experience) and six questions as columns: what happened, why in this order, did it work, how reliable is that judgment, what should have happened, may it be reused. Only verified experience answers all six.

Figure 2. What each representation can answer. A representation is useful for learning only if it can answer the questions a learning claim depends on. Marks are qualitative judgments about the representation as typically recorded, not about any particular system. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative judgment.

The two right-hand columns are where most archives fail. How reliable is the judgment? requires the outcome to name its verifier and that verifier's calibrated error. May it be reused, and for what? requires a rights record attached at capture time. As the companion note on rights as infrastructure argues, a purpose cannot honestly be backfilled onto data captured without one.

Labels have provenance

Not all outcome labels are equal. Recording where data and labels come from is an established documentation practice for datasets (Gebru et al.), and a large audit of dataset licensing and attribution found such information frequently missing or wrong (Longpre et al.). In decreasing order of trust, Ethen's internal architecture work distinguishes:

  1. Ground-truth verification. A deterministic check against acceptance criteria: tests pass, an end state matches, balances reconcile. This is the only tier we consider safe as a training reward without further calibration.
  2. Human judgment. Explicit approval, expert review or rubric grading. Indispensable where no deterministic check exists, but subject to disagreement and to the gap between stated and revealed preference.
  3. Implicit signals. Accept-without-edit, merge, export, no reopen. Useful in aggregate after calibration against the first two tiers; unsafe as a label on a single instance.

Model judges sit between tiers 1 and 2 and need their own calibration. LLM judges show position, verbosity and self-preference biases even when they agree with humans on average (Zheng et al.; Wang et al.; Panickssery et al.). The full treatment is in Evaluating the Evaluators.

The design rule that follows is simple and frequently violated: every label carries its tier and its verifier identity, and a lower-tier label is never silently promoted. A dataset that mixes tiers without recording them cannot be audited, and a model trained on it inherits the noise of its weakest source.

Several anti-patterns recur in agent analytics: treating "run completed" as success; treating a period of user silence as resolution; accepting the agent's own report of success; and counting a positive rating without behavioral corroboration. Each converts the absence of evidence into evidence.

Outcomes are states, not booleans

A second design rule concerns time. Many outcomes are not knowable when a run ends. A code change may pass tests and be reverted a week later. A support ticket may close and reopen. A research report may be accepted and later found to cite a retracted source. Figure 3 shows the lifecycle we propose.

State diagram. Pending leads to provisional success, partial, failure, abandoned or unknown. Provisional success leads to confirmed success when the window closes, or to reopened or reverted. Any terminal state can be superseded by later evidence, which creates a new versioned label.

Figure 3. An outcome label is a state, not a bit. Proposed lifecycle for an outcome label. Success is provisional until a reopen or revert window closes; partial, abandoned and unknown are first-class states; later evidence supersedes rather than overwrites. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (outcome-state semantics).

Three features matter. Success is provisional until a product-specific reopen or revert window closes. Partial success is a first-class verdict, because a task that is 70% done carries information that a binary label discards. Abandoned and unknown are distinct from failure. An abandoned task describes the task, not the agent, and should not train as a negative example without review. And when new evidence arrives, the old label is superseded by a new versioned label rather than overwritten, so that any dataset built from the old label can be identified and rebuilt. The Outcome Warehouse describes how delayed and superseding outcomes are stored.

Failures carry more information per byte

Successful runs of routine tasks are abundant and mostly redundant. Failures, especially failures with an identified cause and a demonstrated fix, are scarce and dense with information. Research on agent failures increasingly focuses on locating the critical step: the earliest point at which an agent had enough information to proceed correctly but did not. TrajDebug, for example, traces each error's lifecycle through a long trajectory to separate errors that caused the final failure from those that were later resolved (Qi et al.). Taxonomies of multi-agent failure (Cemri et al.), automated attribution of which agent failed and when (Zhang et al.), and studies of how agents can learn from their failures (Zhu et al.) point in the same direction. A failure record annotated with its critical step is much closer to verified experience than a thousand unlabeled successes. We develop this in Toward a Failure Genome of Software Agents.

A retention policy based on information value

If labels are the scarce good, retention should follow label value rather than volume. Ethen's internal synthesis proposes scoring each record before long-term retention. The score rises with label tier, with a mined failure and its critical step, with an attached human correction, and with coverage of under-represented task clusters. Two conditions act as gates rather than score components:

  • Rights. A record without a recorded reuse purpose stays within its tenant's boundary and its retention terms, however informative it is.
  • Decontamination. A record derived from an evaluation item never enters a training lineage. Evaluation and training lineages stay separate, enforced with canary strings, temporal cutoffs and deduplication against an exclusion list. Watermarking can make leakage statistically detectable after the fact (Sander et al.).

Under this policy, unlabeled successful runs are sampled for distribution monitoring and otherwise expire. Failures with identified causes are retained preferentially. This is a design proposal; we have not measured its effect on downstream learning, and the weights in any scoring function are judgment calls until tested.

Why telemetry is not a moat

The argument so far implies a sharper statement: agent telemetry is not, by itself, defensible data. Three observations support it. First, trajectory corpora are becoming public. Open environment suites release trajectories and trained verifiers alongside tasks (Pan et al.). Second, a competitor with the same frontier model can generate comparable trajectories on comparable tasks. Third, without outcomes and rights, telemetry cannot be turned into a training or evaluation asset, so its strategic value depends on work not yet done.

What is harder to reproduce is the combination this note describes: outcomes judged by calibrated verifiers, corrections from real reviewers, rare failures with diagnosed causes, and a rights history that permits use. The broader question of which data properties create durability is taken up in What Makes AI Data Defensible?.

Privacy is a property of the representation

Richer representations carry more identifying information. Trajectories can reveal a customer through uncommon error strings, code structure, resource naming, timing and the sequence of tools used. Recent work shows that language models can deanonymize pseudonymous users at scale from unstructured text, reporting up to 68% recall at 90% precision in a closed-world matching setting where classical methods achieved near zero [EXTERNAL PRIMARY-SOURCE RESULT] (Lermen et al.). The lesson for agent data is direct: pseudonymization is not anonymization, and detailed trajectories should be treated as private to their tenant unless a specific re-identification analysis on the actual schema supports a narrower claim. We do not import another paper's attack rate as a property of any particular dataset. We treat it as evidence that the risk is real and must be measured.

Open questions

  • Calibration of implicit signals. How well does accept-without-edit predict verified success, per task family? This must be measured, not assumed.
  • Window lengths. What reopen or revert window balances label latency against label accuracy for each kind of work?
  • Critical-step labeling at scale. Can automated critical-step attribution be trusted on new distributions, and how much expert adjudication does it need?
  • Value of retained successes. What sampling rate of unlabeled successes preserves enough distributional information for monitoring and calibration?

Limitations

This note is conceptual. The six-level taxonomy is a simplification: real systems blur trace and trajectory, and some outcome records already include corrections. The retention policy is a proposal whose weights have not been tested. And the claim that verified experience is more valuable for learning than raw telemetry is itself a hypothesis. Its decisive test is described in How to Test Whether Verified Experience Improves AI Agents.

Conclusion

Telemetry tells a system what it did. Verified experience tells it whether that was right, how confident to be in the judgment, what right would have looked like, and whether it is permitted to learn the lesson. Converting the first into the second takes verifiers with known error, outcome labels that evolve over time, corrections captured at the point of review, and rights recorded at capture. Without those, an archive of agent activity grows without compounding.

FAQ

Is a trace the same as a trajectory? Not quite. A trace is a structured record of execution spans across components. A trajectory is the sequence of observations and actions taken toward a task. A trajectory can be derived from a trace if the trace records the task and the agent's decisions.

Can implicit signals like "accepted without edits" be used as labels? In aggregate, after calibration against stronger labels. On a single instance they are too noisy to use as a training label.

Does verified experience require customer data to leave the customer? No. Verified experience can stay inside a tenant and still improve that tenant's system. Cross-tenant use requires an explicit purpose grant.

References

  1. OpenTelemetry. Semantic conventions for generative AI systems. https://opentelemetry.io/docs/specs/semconv/gen-ai/
  2. Zheng, L. et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685. https://arxiv.org/abs/2306.05685
  3. Wang, P. et al. (2023). Large Language Models are not Fair Evaluators. arXiv:2305.17926. https://arxiv.org/abs/2305.17926
  4. Panickssery, A. et al. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076. https://arxiv.org/abs/2404.13076
  5. Qi, Y. et al. (2026). TrajDebug: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories. arXiv:2608.06346. https://arxiv.org/abs/2608.06346
  6. Cemri, M. et al. (2025). Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657. https://arxiv.org/abs/2503.13657
  7. Zhang, S. et al. (2025). Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. arXiv:2505.00212. https://arxiv.org/abs/2505.00212
  8. Zhu, K. et al. (2025). Where LLM Agents Fail and How They Can Learn From Failures. arXiv:2509.25370. https://arxiv.org/abs/2509.25370
  9. Sander, T. et al. (2025). Detecting Benchmark Contamination Through Watermarking. arXiv:2502.17259. https://arxiv.org/abs/2502.17259
  10. Pan, J. et al. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv:2412.21139. https://arxiv.org/abs/2412.21139
  11. Lermen, S. et al. (2026). Large-scale online deanonymization with LLMs. arXiv:2602.16800. https://arxiv.org/abs/2602.16800
  12. Gebru, T. et al. (2018). Datasheets for Datasets. arXiv:1803.09010. https://arxiv.org/abs/1803.09010
  13. Longpre, S. et al. (2023). The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI. arXiv:2310.16787. https://arxiv.org/abs/2310.16787

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Product

    How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading Paths

    The fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.

  • Product

    Why Long-Running AI Work Needs a Different UX Than Chat

    Long-running AI work needs a different user experience than chat because chat is built on four assumptions that stop being true once work lasts longer than a few minutes: that the work takes seconds, that you are watching, that the result is a single reply, and that you can judge that reply on the spot. AI agents that research, code, operate tools or carry out multi-step tasks now routinely run for minutes or hours, often while the person who started them does something else. That work needs a job with its own identity and state, progress shown as phases rather than a spinner, decision points that reach you at the right moment with the context to answer, pause, resume and recovery that never repeat an action twice, and a result delivered with evidence for review — not a confident final message.

  • Product

    What “Done” Should Mean for an AI Agent

    For an AI agent, "done" should mean that every requirement of the task has been met and that something other than the agent's own report shows it. Precisely: a task is complete when each required obligation is supported by evidence at the level of checking it needs — a passing test, a reconciled record, a confirmed delivery, an approved review — or has been explicitly waived by the person who owns the task. Three refinements make the definition usable. Keep execution success (the steps ran), task success (the outcome was achieved) and business success (it produced value) apart. Treat success as provisional until it can no longer be reversed. And report partial and unknown outcomes as what they are, instead of rounding them up to done.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.