Position Paper · 2026-10-03 · Adaptive Intelligence
Verified Adaptive Intelligence: Learning From Work That Can Be Proven
An AI system should learn from its own work only when that work can be verified, lawfully reused and shown to improve unseen work. Those three conditions are a research program, not a slogan.
Abstract
Deployed AI agents produce a large volume of activity: model calls, tool invocations, approvals, corrections and outcomes. A common belief holds that this activity is a learning signal, so that more usage automatically yields a better system. We argue that the belief is wrong in its naive form and productive in a narrower one. Most agent activity has no trustworthy outcome attached. Much of it cannot legally be reused. And the improvements it seems to support often fail to survive a change of task family, tool version or underlying model. We define verified adaptive intelligence: a system that adapts only from experience whose outcome has been checked by a verifier with known error rates, whose reuse rights are recorded, and whose effect has been measured on held-out work. We describe the mechanism that connects authorized work to verified experience to candidate improvements. We propose an improvement ladder that puts reversible, frozen-weight changes ahead of weight updates. We state five falsifiable hypotheses and name the protocol that tests each one. This paper is an Ethen Research Lab position and agenda. It reports no new measured results.
Why this matters
Foundation models improve mainly through pretraining and post-training done by their developers. Organizations that deploy agents on top of those models face a different question: how should their system get better at their work between model releases, and how can they tell when it has?
The usual answer is a data flywheel: more usage produces more data, more data produces a better model, and a better model attracts more usage. Each link in that chain has been weakened by published evidence. Training on recursively generated content can degrade a model's coverage of the tails of a distribution [EXTERNAL PRIMARY-SOURCE RESULT] (Shumailov et al.). The size and conditions of that effect are still debated (Seddik et al.; Borji). Scaling preference-based post-training shows diminishing returns along several axes (Hou et al.). Proxy rewards can be optimized in ways that diverge from the intended objective, and this follows from how reward hacking is defined, not from carelessness (Skalse et al.; Gao et al.). None of this makes learning from deployment impossible. It does mean that the burden of proof sits with any claim that a deployment is learning.
The burden matters because agents now act. When a system edits code, files tickets, moves records or spends money, its errors are effects in the world, not just bad text. An organization that cannot say what its agents did, whether it worked, and whether last month's "improvement" holds up on next month's work cannot responsibly widen their autonomy.
Defining the terms
We use four terms precisely.
Experience is the structured record of a unit of work. It covers what was intended, under whose authority, which actions were taken, what effects were observed, how the outcome was judged, which corrections occurred and what it cost. A trace that records only activity ("called API X, produced text Y") is not yet experience in this sense. The companion note From AI Traces to Verified Experience develops the distinction.
Verification is a judgment about whether the work met its acceptance criteria, made by a mechanism whose error rates are known. A passing test suite is a verifier. So is a reconciliation of account balances, or an expert scoring against a rubric. An agent reporting its own success is not.
Adaptation is any change to the system's behavior that is caused by experience. That includes routing rules, context policies, recovery procedures, versioned skills, verifier thresholds and, last, model weights.
Transfer is improvement that persists when something important changes: a new task family, a new tool version, or a new underlying model. An adaptation that helps only on the tasks it was fitted to is memorization with extra steps.
Verified adaptive intelligence, then, is the property of a system that (1) adapts only from experience whose outcome has been verified with known error, (2) knows the rights that govern reuse of that experience, and (3) measures whether each adaptation transfers before relying on it.
Why raw usage is insufficient
Figure 1 shows the gates a run must pass before it can support a learning claim. Each gate removes a large class of activity, and each corresponds to a distinct failure of the naive flywheel.
Figure 1. Most activity never becomes learning signal. Each gate removes runs that cannot support a learning claim. Only experience that passes all four gates, and then improves held-out work, counts as capability. The gates are a design proposal; no Ethen volumes are implied. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis (research hypothesis); not measured.
Absent labels. Most agent runs end without a trustworthy outcome. The model returned text, the tool returned 200, the user closed the tab. None of these establishes that the task succeeded. Treating "run completed" as success is the most common error in agent analytics.
Biased labels. Where outcomes exist, they are not missing at random. Easy tasks get verified more often than hard ones. Tasks routed to a particular model are observed only under that model. Filtering to successes biases a corpus toward the work the system already does well. Selection bias in routing data is serious enough that we treat it in a dedicated note on propensity logging.
Gamed labels. Any proxy that becomes a training target invites optimization against the proxy. A verifier that accepts 5% of bad work looks harmless as a dashboard metric. As a reward, it is an instruction to find that 5%. The bound this places on learning is the subject of Evaluating the Evaluators: Reward Integrity for AI Agents.
Unusable labels. Experience produced inside a customer's boundary usually belongs to that customer's purposes. A record that cannot lawfully enter a training or evaluation build is not a learning asset, however informative it is.
A fifth problem concerns defensibility more than learning. Generic trajectories are becoming abundant. Open environment suites now release trajectories and trained verifiers alongside their tasks (Pan et al., SWE-Gym). When trajectories are cheap, the scarce input is the environment, the verifier and the outcome label, not the log.
What existing research shows
Several research lines show that experience can improve agents without retraining a foundation model. Reflexion stores verbal self-critiques between attempts (Shinn et al.). Voyager accumulates an executable skill library in an open-ended environment (Wang et al.). ExpeL extracts insights from past successes and failures (Zhao et al.). Agent Workflow Memory induces reusable workflows from past trajectories (Wang et al.). SkillWeaver lets web agents discover and refine skills that other agents can reuse (Zheng et al.). On the training side, reinforcement learning in a high-fidelity simulation of a customer-support organization improved held-out task pass rates. The gains also carried over to out-of-distribution tool-use benchmarks [EXTERNAL PRIMARY-SOURCE RESULT] (Mehta et al., CoreCraft). Process-level verification can supervise reasoning more effectively than outcome-only feedback in mathematics (Lightman et al.).
These results are encouraging, and they leave three gaps that matter for deployed work.
First, verification is usually assumed rather than measured. Benchmarks ship with graders, and papers rarely report those graders' false-accept rates on the distribution being learned from.
Second, rights are outside the loop. Research corpora are assembled once. Deployed experience arrives continuously under contractual purposes that can expire or be revoked.
Third, transfer across model generations is rarely the target. A skill library that helps one model may be useless or harmful for its successor. Prompt-level behavior is sensitive to formatting choices that carry no meaning (Sclar et al.), and hosted model behavior can drift between versions (Chen et al.).
Measurement of capability growth itself is improving. Task-horizon methods track how long a task an agent can complete at a given success rate (Kwa et al.), and cost-aware evaluation has been argued for explicitly (Kapoor et al.). We build on both.
The Ethen thesis
Ethen research hypothesis. Verified experience, represented with explicit commitments and rights metadata, can improve routing, recovery, context selection and skills more than an equal-cost volume of unverified experience. Part of that improvement survives a model upgrade.
Two clarifications keep the hypothesis honest. It is comparative: the baseline is not "no data" but "the same budget spent on unverified data or on simple rules". And it is partial: we expect some adaptations to be erased by a better model, and we intend to measure which ones.
The thesis also implies an architecture. Work begins under a mandate that states what is authorized. A commitment graph records what the work is obligated to achieve and what remains unfinished. Execution produces a work receipt: a single record of authority, actions, effects, verification, cost and rights. Receipts that pass the gates in Figure 1 become verified experience. Candidate changes are fitted to it and then tested on sealed, held-out work before any promotion.
Where labels come from
The quality of what a system can learn is capped by the quality of its labels. Figure 2 orders label sources by trust, following the verification order proposed in Ethen's internal architecture work: deterministic invariants first, then programmatic verifiers, expert rubrics, calibrated model judges and, last, implicit behavioral signals.
Figure 2. Where outcome labels come from, and what each can support. Label sources ordered by trust. The hierarchy follows the outcome-verification order proposed in Ethen's internal architecture work (deterministic invariant before programmatic verifier, expert rubric, calibrated judge and implicit signal). Marks are qualitative judgments, not measurements. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative judgment.
A simple identity shows why the order matters. Let a task class have true success rate p. Let the verifier have false-accept rate f and false-reject rate g. The observed success rate is then
p<sub>obs</sub> = p(1 − g) + (1 − p)f.
Any policy that maximizes p<sub>obs</sub> can raise it either by raising p or by moving into regions where f is high. Only the first is improvement. Unless f is measured and small where the policy goes, a learning curve built on p<sub>obs</sub> can rise while real capability stays flat. This is why the program puts verifier measurement before learning. The protocol for that measurement is in How Should We Measure the Reliability of LLM Verifiers?.
An improvement ladder
Not all adaptations carry the same risk. Figure 3 orders them by cost, reversibility and attributability.
Figure 3. An improvement ladder with a reward-integrity gate. Proposed order for adapting a deployed system: cheap, reversible changes first; weight changes last, offline, and only after the reward-integrity gate is green. The ordering is an Ethen research hypothesis (H2) that the program tests rather than assumes. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis (learning ladder); hypothesis.
The first three rungs leave model weights frozen. Routing and configuration choices, harness and context policies, and versioned skills can be changed in hours, rolled back in seconds and attributed with ablations. The fourth rung is a gate rather than an adaptation: no reward-based training proceeds until the verifiers that would supply its rewards have published error rates and have survived deliberate attempts to game them. Policy optimization and weight updates come last. They run offline and in batches, are evaluated on sealed sets, and are promoted only through the same gates as every other change.
This ordering is itself a hypothesis (H2 below). It is supported by an internal synthesis of practitioner evidence suggesting that harness- and routing-level gains are common and inexpensive. Whether that holds for Ethen's workloads is unknown.
Relation to the pursuit of general intelligence
We avoid using general intelligence as a label, but we do not avoid the question. The approach here is one tractable sub-question within it: under what conditions does experience become capability that generalizes? Digital work is a useful setting for studying that question. Goals are explicit. Effects can be observed. Failures are frequent enough to study. And environments can be reset, branched and replayed.
Three longer-term research directions follow from the thesis. All three are speculative and none is under way:
- Compositional experience learning. Can capabilities learned separately be combined to complete unseen task families, compared against explicit symbolic composition and retrieval at equal resources?
- Uncertainty-aware state models. Can a learned model of how digital state changes under actions, including its own uncertainty, improve planning and the decision to stop?
- Long-trajectory credit assignment. Which decisions in a long run actually caused its eventual success, and can verified intermediate goals or controlled interventions identify them better than a single terminal reward?
If those directions produce replicated cross-domain gains, they would contribute to learning science, not merely to integration. If they do not, the operational program still stands on its own.
A falsifiable agenda
Figure 4 states five hypotheses, each with a refutation condition and a protocol paper in this library.
Figure 4. Five falsifiable hypotheses and what would refute them. The agenda is stated so that it can fail. Each hypothesis maps to a protocol paper in this library. None has been tested yet. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen research hypotheses; untested.
- H1 — Verification value. At matched curation cost, verified examples produce larger held-out gains than unverified ones. Tested in How to Test Whether Verified Experience Improves AI Agents.
- H2 — Ladder ordering. Most early gains come from the frozen-weight rungs.
- H3 — Transfer. Some fraction of the gains survives a model upgrade. The fraction is measured with a capability transfer ledger.
- H4 — Recovery. Counterfactual recovery examples generalize to failure mechanisms held out of training. See Recovery Atlas.
- H5 — Economics. Cost per verified outcome falls relative to strong static rules.
The benchmark umbrella for these tests is Ethen VerifiedWork. Every experiment in the program follows one contract: a falsifiable hypothesis; the strongest practical baseline; pinned versions of every dataset, environment, grader and model; rights-cleared development data; a sealed holdout managed by someone other than the experiment's author; a statistical plan written before the run; ablations; full cost accounting; and publication of negative results.
Competing explanations
An observed improvement can have causes other than verified experience, and the protocols are designed to separate them:
- More compute. The adapted system may simply use more tokens, retries or verification. Comparisons are made at matched budget, and cost is reported as cost per verified outcome.
- Infrastructure effects. Caching and provider changes can lower cost without any change in decision quality. Routing studies factor these out explicitly.
- Contamination. Held-out tasks can leak into prompts, skills or training builds. Sealed registries, canary strings and temporal splits guard against this.
- Model drift. A provider's model can improve or regress between measurements. All comparisons pin model versions and run paired designs.
- Grader exploitation. Improvement on the grader can mask no improvement on the task. Grader audits run alongside each comparison.
Limitations
This paper proposes a program. It does not report results. Three limitations are fundamental.
First, verification has a ceiling. Many valuable tasks, such as research synthesis, design and negotiation, resist deterministic verification. For those tasks the program depends on expert judgment, which is expensive and imperfect.
Second, rights may constrain learning more than technology does. Customers can and often will prohibit reuse of their data. A program that only works with broad reuse grants would be fragile, so the design must deliver value when experience stays inside one tenant.
Third, transfer may be small. If adaptations rarely survive a model upgrade, the long-run value shifts from accumulated capability toward measurement and assurance. That shift is analyzed in Why Better Foundation Models May Make Evaluation More Valuable, Not Less.
Conclusion
Learning from deployment is real, but it is narrow. It depends on knowing what happened, whether it worked, whether the record may be used, and whether the lesson holds somewhere new. Verified adaptive intelligence treats each of those as a measurement problem with an explicit error rate, rather than as an assumption. The research question we commit to is empirical and can fail: does verified experience produce capability that transfers? The rest of this library describes how we intend to find out.
FAQ
Is verified adaptive intelligence the same as continual learning? No. Continual learning usually refers to updating model weights over time without forgetting. Verified adaptive intelligence covers a broader set of adaptations, most of which leave weights frozen. It adds verification, rights and transfer requirements that continual-learning research usually does not impose.
Does this mean training models on customer data? Not by default. The design assumes that customer experience stays inside the customer's boundary unless an explicit, revocable purpose grant says otherwise. See Rights as Infrastructure.
Why not just wait for better foundation models? Better models raise the baseline, which every protocol here uses as the comparison. They do not remove the need to know what an agent did, whether it worked, and whether a change holds up on new work.
Has Ethen demonstrated these hypotheses? No. All five hypotheses are untested. The protocol papers describe the experiments.
Related research
- From AI Traces to Verified Experience — defines the unit of verified experience.
- Work Receipts: A Verifiable Record for Autonomous AI Work — the record that makes work provable.
- Commitment Graphs: Why AI Agents Need to Know What Is Still Unfinished — obligation semantics behind 'finished'.
- Recovery Atlas: Teaching AI Agents When to Retry, Reconcile, Escalate, or Stop — recovery as a learnable capability.
- Evaluating the Evaluators: Reward Integrity for AI Agents — verifier error bounds what can be learned.
- The Capability Transfer Ledger: Measuring Whether AI Skills Survive Model Upgrades — measuring transfer across model generations.
- Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action — benchmark umbrella for the agenda.
- How to Test Whether Verified Experience Improves AI Agents — the decisive experiment for the thesis.
References
- Shumailov, I. et al. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493. Published as "AI models collapse when trained on recursively generated data", Nature 631 (2024). https://arxiv.org/abs/2305.17493
- Seddik, M. E. A. et al. (2024). How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse. arXiv:2404.05090. https://arxiv.org/abs/2404.05090
- Borji, A. (2024). A Note on Shumailov et al. (2024). arXiv:2410.12954. https://arxiv.org/abs/2410.12954
- Hou, Z. et al. (2024). Does RLHF Scale? Exploring the Impacts From Data, Model, and Method. arXiv:2412.06000. https://arxiv.org/abs/2412.06000
- Skalse, J. et al. (2022). Defining and Characterizing Reward Hacking. NeurIPS 2022. arXiv:2209.13085. https://arxiv.org/abs/2209.13085
- Gao, L. et al. (2022). Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760. https://arxiv.org/abs/2210.10760
- Pan, J. et al. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv:2412.21139. https://arxiv.org/abs/2412.21139
- Shinn, N. et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. https://arxiv.org/abs/2303.11366
- Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
- Zhao, A. et al. (2023). ExpeL: LLM Agents Are Experiential Learners. arXiv:2308.10144. https://arxiv.org/abs/2308.10144
- Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
- Zheng, B. et al. (2025). SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills. arXiv:2504.07079. https://arxiv.org/abs/2504.07079
- Mehta, S. et al. (2026). EnterpriseBench CoreCraft: Training Generalizable Agents on High-Fidelity RL Environments. arXiv:2602.16179. https://arxiv.org/abs/2602.16179
- Lightman, H. et al. (2023). Let's Verify Step by Step. arXiv:2305.20050. https://arxiv.org/abs/2305.20050
- Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
- Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
- Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499
- Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- From AI Traces to Verified Experience
Logs, traces, trajectories, outcomes and corrections are not the same asset. A research note on what turns agent telemetry into verified experience.
- What Makes AI Data Defensible?
A position paper on the AI data moat: why volume is not defensibility, and eight axes, from rights to outcome density and transfer, that decide what compounds.
- The Outcome Warehouse: Turning Completed AI Work Into Research Assets
A research note on the Outcome Warehouse: storing verified, versioned AI outcome data with corrections, cost, lineage, rights and delayed business results.
Explained on the Ethen Blog
- Why Ethen Is a Family of Specialized AI Apps, Not One App
Ethen is organized as a family of focused apps rather than one universal chat window because different kinds of AI work need different things to persist, different controls, and different evidence. A quick question needs a thread. A research investigation needs sources and branches that survive across sessions. A code change needs a repository, tests and review. A goal-driven job needs a plan, approvals and a record of what actually happened. Ethen gives each kind of work one canonical home — Chat, Code, Studio, Research, Designer, Founder, the Platform products and Desktop — and keeps the account, model access, permissions and evidence shared underneath, so it still works as one Ethen.
- What We're Building Across Ethen: October 2026 Update
As of October 2026, Ethen's work falls into five areas. We are organizing Ethen into focused apps that share one foundation. We are hardening that foundation for AI work that runs for minutes or hours: durable jobs, completion that depends on evidence, honest handling of unknown outcomes, and approvals tied to specific actions. We are building a model knowledge layer so people can choose models with sources rather than guesses. We are extending Ethen to local, on-device AI through Desktop. And Ethen Research Lab now publishes its research in public, with every paper labeled by evidence status. This update separates what our public posts describe as implemented from what is design direction, and it makes no launch or date commitments.
- Why Ethen Research Lab Publishes Its Work in Public
Ethen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.