Research Protocol · 2026-10-03 · Adaptive Intelligence
How to Test Whether Verified Experience Improves AI Agents
The central bet of the Verified Adaptive Intelligence agenda is that experience with verified outcomes is worth more than experience without them. This protocol describes the experiment that could show the bet is wrong.
Abstract
Many agent systems assume that recording their own work will make them better: more trajectories, more data, more capability. The research program behind this series makes a narrower and testable claim about learning from experience for AI agents. Experience that is rights-cleared, verified, linked to outcomes and rich in failures, corrections and recoveries may produce more transferable improvement than raw traces or volume-matched unverified data. This protocol specifies how to test that claim. Five data conditions are compared at matched volume: raw trajectories, successful-only trajectories, verified outcomes, verified failures with their recoveries, and rights-aware value-selected experience. Each is applied through a ladder of interventions that starts with the cheapest and least risky, namely retrieval, skill updates, routing and context policy, and reaches narrow model training only if cheaper interventions show signal. Outcomes are measured on a pinned base model, on held-out task families, and again after a model or tool version change. The primary result is a set of data gain curves: verified success as a function of data volume, per condition. A reward-integrity gate prevents verifier errors from being learned as truth. No measurements exist. The thesis under test is described in Verified Adaptive Intelligence; a working title that implied results was replaced because none have been measured.
The claim, stated so it can fail
The weakest version of the thesis is that experience helps. That version is nearly unfalsifiable and not very interesting: almost any additional relevant data helps something somewhere. The version worth testing is comparative and conditional:
At matched data volume and matched intervention cost, experience that is verified, outcome-linked and enriched for failures and recoveries yields larger and more durable gains on held-out work than raw or success-filtered trajectories.
Each qualifier is load-bearing. Matched volume rules out the explanation that curated data wins only because there is more of it. Matched cost rules out the explanation that it wins only because more effort went into using it. Held-out work rules out memorization. Durable means the gain survives a model or tool version change, which separates lasting capability from compensation for a particular model's weaknesses.
What counts as verified experience, and the different representations it can take, is defined in From AI Traces to Verified Experience. The argument that raw data is not automatically an asset, and that training utility is the decisive property of a dataset, is developed in What Makes AI Data Defensible?. This protocol turns that argument into an experiment.
What prior work suggests
Several lines of research make the hypothesis plausible without settling it.
Experience reuse works on familiar tasks. Agents that store verbal reflections on failed attempts improve on subsequent attempts (Shinn et al.). Agents that extract insights and examples from collections of past experiences improve on related tasks (Zhao et al.). Workflows induced from past trajectories improve web agents on later tasks (Wang et al., 2024), and growing skill libraries support continued progress in open-ended environments (Wang et al., 2023).
Training on environments with checkable outcomes can transfer. Training software agents on real repositories with executable tests produced substantial gains on held-out benchmarks, and verifiers trained on sampled trajectories improved results further (Pan et al.). Training on a high-fidelity enterprise simulation with expert-authored rubrics produced gains on held-out tasks and on out-of-distribution benchmarks (Mehta et al.). Process-level supervision outperformed outcome-only supervision for training reward models on mathematical reasoning (Lightman et al.).
Volume alone has limits. Post-training methods show diminishing returns as data and compute grow (Hou et al.), and training on model-generated content can erode the tails of the original distribution (Shumailov et al.; Seddik et al.). Agent trajectories are partly model-generated, so naive accumulation is not obviously beneficial.
Rewards can be gamed. Optimizing against an imperfect reward model eventually degrades true performance (Gao et al.), and reward misspecification produces systematic misbehavior (Pan et al., 2022; Skalse et al.). Verified experience is only as good as its verifiers.
None of these studies compares verified, failure-rich experience with volume-matched alternatives on held-out agent work and across a model change. That is the gap.
Five data conditions
Figure 1 shows the conditions crossed with the interventions.
Figure 1. Five data conditions across the intervention ladder. Every data condition is volume-matched and applied through interventions in order of cost and risk. Higher rungs are attempted only when lower rungs show signal; training rungs run only in sealed environments on rights-cleared data with the base model pinned. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
- Raw trajectories. Complete traces of agent work as logged, with no outcome labels beyond whether the run ended.
- Successful-only trajectories. Raw trajectories filtered to runs that completed, using the runtime's own success signal. This is the common practice of keeping "good" examples.
- Verified outcomes. Trajectories labeled by verifiers that passed the reward-integrity gate, including partial and failed outcomes, with the label's provenance attached.
- Verified failures and recoveries. Verified trajectories enriched for failures, the critical step at which they went wrong, the correction applied and the recovery outcome, drawn from sources such as the Recovery Atlas and described with the axes of the Failure Genome.
- Rights-aware, value-selected experience. Verified experience selected for expected training value, such as novelty, label strength and failure informativeness, and restricted to records whose rights permit this purpose.
All conditions are drawn from the same underlying task pool, stored as outcome-linked records in the Outcome Warehouse, and are volume-matched in two ways, by number of episodes and by number of tokens. Results are reported under both, because conditions differ in average trajectory length.
The intervention ladder
Interventions are tried in order of cost and risk, a sequence we call the learning ladder:
- Retrieval. Relevant past experience is retrieved into context at task time.
- Skill update. Experience is distilled into revised skills or procedures, validated as described for Skill IR.
- Routing policy. Experience informs which execution configuration is chosen for which task.
- Context policy. Experience informs what context is compiled for which kind of work.
- Narrow model training. A small model, such as a verifier, a router or a recovery classifier, is trained on the data.
- Post-training of the agent model. Fine-tuning or reinforcement learning of the agent model itself.
A higher rung is attempted only if a lower one shows signal for some data condition, or if there is a specific reason to expect the lower rung cannot express the gain. Rungs 5 and 6 are run only in sealed environments, on rights-cleared data, with the base model pinned. This ordering reflects a design decision to keep weights frozen until cheaper adaptation has been exhausted. It also means the study may conclude before reaching training at all, which would itself be informative.
Gates before any learning
Figure 2 shows the gates every data condition must pass.
Figure 2. Gates every data condition must pass. No record reaches a learner unless its rights permit the purpose, its verifier passed the reward-integrity gate, and it is excluded from sealed evaluation families. Training and evaluation verifiers are kept separate. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Rights gate. Every record used must carry rights that permit the stated research purpose. Datasets are compiled by the procedure in A Rights-Aware Dataset Compiler, which fails closed on unknown rights and produces a signed manifest. Records whose rights are withdrawn later are traceable through lineage, and any affected result is identified.
Reward-integrity gate. Labels from a verifier are used only if the verifier's measured false-positive and false-negative rates on that task family are within pre-registered limits, measured as described in How Should We Measure the Reliability of LLM Verifiers?. The broader argument is in Evaluating the Evaluators. A verifier that accepts wrong answers would turn condition 3 into a dataset of confidently mislabeled examples.
Separation of training and evaluation verifiers. The verifier that labels training data is not the one that scores evaluation. Where possible, evaluation uses deterministic checks; where a judge is needed, it comes from a different model family. Without this separation, a learner can improve on the verifier's quirks rather than on the task, and the gain curve would measure that.
Contamination controls. Evaluation families are sealed, marked with canaries and excluded from all training conditions by near-duplicate filtering, with temporal splits where possible.
Evaluation: pinned, held out and re-tested after change
Figure 3 shows the evaluation structure.
Figure 3. Evaluation: pinned, held out, then re-tested after change. Gain curves are measured on a pinned base model, on families never used in training. The best intervention per condition is then re-tested after a model upgrade and a tool version change to separate durable capability from compensation for one model's weaknesses. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Pinned base model. All gain curves are measured on a fixed base model version, so that differences reflect data, not model drift.
Held-out families. Evaluation uses task families absent from all training conditions, plus held-out instances of training families. A condition that wins only on familiar families has shown memorization or narrow fitting, not transferable capability.
Version change. The best-performing intervention for each condition is re-evaluated after a model upgrade and after a tool version change. A gain that disappears under a new model was compensating for the old model's weakness. A gain that persists is the property the thesis predicts. The transfer measurements follow the Capability Transfer Ledger.
The data gain curve
For each condition and intervention, data volume is increased by doubling, and verified success and cost per verified outcome are measured at each step on the sealed evaluation set. The resulting curve is the primary result. Three summaries are pre-registered: the gain at the largest common volume, the slope per doubling over the upper half of the range, and the volume at which the curve flattens.
The primary contrast is between condition 4 or 5 and condition 2 at matched volume, because success-filtered trajectories are the strongest common baseline. Condition 1 provides a floor, and condition 3 isolates the value of verification from the value of failure enrichment.
Hypotheses and what would falsify them
Pre-registered as [PROPOSED TARGET]s:
- H1. On held-out families, condition 4 exceeds condition 2 in verified success at matched volume for at least one intervention below model training.
- H2. The slope of the gain curve per doubling is steeper for condition 5 than for condition 1.
- H3. At least half of the H1 advantage persists after a model version change.
The thesis is weakened if conditions 3 to 5 do not outperform condition 2 at matched volume on held-out families for any intervention. It is strongly weakened if they do so on familiar families only, or only before a model change. The study is designed so that both outcomes are publishable. Figure 4 states in advance how each result will be read.
Figure 4. What each result would mean. Outcomes are stated before the study so that a negative result cannot be reinterpreted afterwards. A flat curve for verified experience would move the program's emphasis from learning to assurance. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed); pre-registered interpretation.
A flat curve for verified experience would mean that the value of verification lies in assurance and accountability rather than in learning, which would change where this program invests.
Competing explanations
- Curation effort. Conditions 4 and 5 involve more processing. Intervention cost, including human and compute effort, is matched or reported.
- Difficulty mix. Failure-enriched data over-represents hard tasks. Difficulty is measured per task and controlled in analysis.
- Label leakage. If verifier signals appear in the training data in a form the learner can exploit at evaluation time, gains are spurious. The verifier separation above addresses this.
- Base-model ceiling. If the pinned model is already near ceiling on some families, no condition can show gains there. Such families are reported as saturated.
- Selection by the designers. Condition 5's value-selection rule is fixed before training and not tuned on evaluation results.
Statistical analysis
Gain curves are estimated per family with intervals from a cluster bootstrap over tasks and repeated training seeds, because training variance can rival the effects of interest. Contrasts between conditions at matched volume are paired by evaluation task. Multiplicity across conditions, interventions and families is controlled for false discovery (Benjamini & Hochberg), with H1 confirmatory and the rest secondary. Every run is recorded in an experiment registry with data manifests, model versions, verifier versions and seeds, so that any result can be reproduced or withdrawn if the rights of its data change.
Limitations
The protocol has not been run, and its thresholds are proposed. Building five volume-matched conditions requires a large pool of rights-cleared, verified work, which may take considerable time to accumulate; early runs may be limited to a few families with strong verifiers, such as software tasks with tests. Results will depend on the base model, the intervention recipes and the families studied, and may not generalize to other domains. The value-selection rule in condition 5 is one choice among many. Training interventions raise legal questions about the use of customer-derived data that must be resolved before those rungs are run. Finally, Ethen has an interest in a favorable outcome; independent replication on public environments is part of the plan, not an afterthought.
Conclusion
The thesis that verified experience makes agents better is either a real scientific finding or a convenient story. The difference is whether it survives a comparison at matched volume, on held-out work, after a model change, against strong alternatives. This protocol specifies that comparison and states in advance what result would weaken the thesis. If verified, failure-rich experience produces steeper and more durable gain curves, the program has a foundation. If it does not, the program should say so and move its effort to where the evidence points.
FAQ
Why not train a model on all the data and see what happens? Because that would not distinguish the value of verification from the value of volume, effort or familiarity. Matched conditions and held-out families are the point.
Why start with retrieval instead of training? It is cheaper, reversible and easier to attribute. Training is attempted only if cheaper interventions show signal.
What would count as a negative result? Verified, failure-rich experience failing to beat success-filtered trajectories at matched volume on held-out families, or beating them only before a model change.
Related research
- Verified Adaptive Intelligence: Learning From Work That Can Be Proven — the thesis under test.
- From AI Traces to Verified Experience — what counts as verified experience.
- Evaluating the Evaluators: Reward Integrity for AI Agents — reward integrity gate.
- What Makes AI Data Defensible? — training utility as a data axis.
- A Rights-Aware Dataset Compiler for AI Training and Evaluation — rights-cleared training builds.
- How Should We Measure the Reliability of LLM Verifiers? — verifier measurement protocol.
References
- Shinn, N. et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. https://arxiv.org/abs/2303.11366
- Zhao, A. et al. (2023). ExpeL: LLM Agents Are Experiential Learners. arXiv:2308.10144. https://arxiv.org/abs/2308.10144
- Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
- Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
- Pan, J. et al. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv:2412.21139. https://arxiv.org/abs/2412.21139
- Mehta, S. et al. (2026). EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments. arXiv:2602.16179. https://arxiv.org/abs/2602.16179
- Lightman, H. et al. (2023). Let's Verify Step by Step. arXiv:2305.20050. https://arxiv.org/abs/2305.20050
- Hou, Z. et al. (2024). Does RLHF Scale? Exploring the Impacts From Data, Model, and Method. arXiv:2412.06000. https://arxiv.org/abs/2412.06000
- Shumailov, I. et al. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493. https://arxiv.org/abs/2305.17493
- Seddik, M. E. A. et al. (2024). How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse. arXiv:2404.05090. https://arxiv.org/abs/2404.05090
- Gao, L. et al. (2022). Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760. https://arxiv.org/abs/2210.10760
- Pan, A. et al. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. arXiv:2201.03544. https://arxiv.org/abs/2201.03544
- Skalse, J. et al. (2022). Defining and Characterizing Reward Hacking. arXiv:2209.13085. https://arxiv.org/abs/2209.13085
- Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Verified Adaptive Intelligence: Learning From Work That Can Be Proven
A research agenda for AI agents that learn only from experience that is verified, rights-cleared and shown to transfer across tasks, tools and model generations.
- From AI Traces to Verified Experience
Logs, traces, trajectories, outcomes and corrections are not the same asset. A research note on what turns agent telemetry into verified experience.
- What Makes AI Data Defensible?
A position paper on the AI data moat: why volume is not defensibility, and eight axes, from rights to outcome density and transfer, that decide what compounds.
Explained on the Ethen Blog
- Why Ethen Research Lab Publishes Its Work in Public
Ethen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.