Skip to content

EthenEthenEthen

Methods Paper · 2026-10-03 · Adaptive Intelligence

The Capability Transfer Ledger: Measuring Whether AI Skills Survive Model Upgrades

Publication type
Methods Paper
Research program
Adaptive Intelligence
Published
Authors
Ethen Research Lab
Reading time
12 min read

"This skill is portable" is a claim. "This skill transferred to these model and tool versions, degraded on that one, and made this one worse" is a measurement. We propose a ledger for the second.

Cover image for "The Capability Transfer Ledger: Measuring Whether AI Skills Survive Model Upgrades". Decorative abstract motif; contains no data.

Abstract

Agent systems invest heavily in skills, procedures and configurations that improve performance on particular work. Each investment rests on an assumption that its value will persist when the underlying model, tools or task mix change. That assumption is rarely tested. When it fails, skills can degrade silently, or worse, make a new model perform below its own baseline. This methods paper proposes the Capability Transfer Ledger: an accumulating record of paired measurements of each skill version under combinations of model version, tool version and task family, always with and without the skill. We define four estimands: source gain, target gain, retention and negative transfer. We also define a fifth condition, obsolescence, in which a skill stops helping because the new model no longer needs it. We describe paired designs, reliability across repeated trials, multiple-comparison control and sparse sampling across a large combination space, together with the status rules that turn measurements into compatibility decisions. Capability transfer is the central scientific question of the Verified Adaptive Intelligence agenda. This paper proposes how to measure it and reports no measurements.

Why transfer must be measured

The research case for experience-driven agents rests on skills and memories that accumulate and are reused (Wang et al., Voyager; Wang et al., Agent Workflow Memory; Zheng et al., SkillWeaver). Deployed systems, however, rarely keep the same model for long. Providers release new versions; organizations switch providers for price or capability; tool APIs change. Three findings suggest that accumulated skills may not survive these changes intact. Hosted model behavior can shift substantially between versions (Chen et al.). Prompt performance is sensitive to formatting choices that carry no meaning (Sclar et al.). And model updates in other settings introduce instance-level regressions, called negative flips, even when aggregate accuracy improves (Yan et al.; Srivastava et al.; Echterhoff et al.).

The research agenda in Verified Adaptive Intelligence asks whether verified experience produces capability that transfers. If transfer is near zero, accumulated skills are a recurring cost of re-learning, not an accumulating asset. If transfer is substantial but uneven, an organization needs to know which skills to trust after each change. Either way, the answer has to come from measurement.

The ledger

The ledger is an append-only collection of entries (Figure 1). Each entry is a measurement of one skill version under one condition. A condition is a combination of model version, tool interface versions and task family.

Four input dimensions on the left: skill and version, model and version, tool interface versions, task family. They combine into a ledger entry in the middle containing: success with skill, success without skill (paired baseline), transfer gain with confidence interval, pass^k reliability, cost per verified outcome, sample size, verifier version, date. On the right, the entry's status: validated, degraded, negative transfer, or untested.

Figure 1. What one ledger entry records. Each entry is one measurement of one skill version under one combination of model version, tool version and task family, with and without the skill, on paired tasks. The ledger is the accumulation of such entries. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen methods proposal.

Each entry records the verified success rate with the skill and without it, measured on the same paired tasks. It also records the difference and its confidence interval, reliability across repeated trials, cost per verified outcome, the sample size, the verifier version used and the date. From these, a status is derived: validated, degraded, negative transfer, obsolete or untested.

The representation of skills being measured is described in Skill IR. The ledger does not depend on that representation, though. It can measure prompt recipes, workflow memories, routing policies or recovery policies. What it requires is that each skill can be switched on and off on the same task.

Estimands

For each skill, we define the quantities of interest using a two-by-two design (Figure 2).

Two-by-two grid: rows are source condition and target condition; columns are without skill and with skill. Arrows show gain in source equals with minus without in source; gain in target equals with minus without in target. Below: retention equals target gain divided by source gain; negative transfer when target gain is reliably below zero; obsolescence when target gain is near zero because the baseline already succeeds.

Figure 2. Four estimands from a two-by-two design. For each skill, measure success with and without the skill in the source condition (where it was developed) and the target condition (new model, tool or family), on paired tasks. Gains in each condition define transfer, retention and negative transfer. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen methods proposal.

Let S₀ and S₁ be the verified success rates without and with the skill in the source condition, where the skill was developed. Let T₀ and T₁ be the corresponding rates in a target condition. Then:

  • Source gain is S₁ − S₀: the skill's value where it was built.
  • Target gain is T₁ − T₀: the skill's value in the new condition. This is the primary estimand.
  • Retention is the ratio of target gain to source gain: the share of the skill's value that survived. Ratios are unstable when source gain is small, so retention should be reported only when source gain is reliably positive, and with an interval.
  • Negative transfer is a target gain reliably below zero: the skill makes the new condition worse than having no skill.

A fifth condition deserves its own name. Obsolescence occurs when target gain is near zero because the baseline T₀ is already high: the new model succeeds without help. Obsolescence is not a failure of the skill; it is a success of the model. It should be recorded separately from degradation, because the right response differs. An obsolete skill can be retired. A degraded one may need a new realization.

Measuring T₀ is essential and easy to forget. Without it, a high T₁ after a model upgrade looks like successful transfer when the skill may be contributing nothing, or even holding the model back.

Design

Paired tasks. Every comparison runs with and without the skill on the same task instances, from the same initial state, with the same tools, budget and verifier. Pairing removes task difficulty from the comparison and sharply reduces the number of tasks needed. Binary paired outcomes can be analyzed with McNemar's test or with intervals for paired differences.

Repeated trials. Agents are stochastic. Each task should be run several times per arm, and results reported both as mean success and as reliability: the probability of succeeding on all of k trials, the pass^k measure introduced with τ-bench (Yao et al.). A skill can improve average success while reducing reliability, and that trade matters in production.

Held-out families. Transfer to a new task family is the hardest and most interesting case. Families used to develop a skill must be excluded when measuring family transfer. Otherwise the measurement reports memorization.

Clustering. Tasks from the same template, customer or source are correlated. Intervals should be computed with clustering by family, or with paired bootstrap resampling at the family level.

Multiple comparisons. A ledger with many skills and conditions will produce some apparently significant negative transfers by chance. Status decisions should control the false discovery rate across the comparisons made at each review (Benjamini & Hochberg). The policy should be pre-registered rather than chosen after seeing results.

Verifier versioning. A change of verifier between source and target measurements can masquerade as transfer or its absence. Entries record verifier versions, and comparisons should hold the verifier fixed. Verifier reliability itself is covered in Evaluating the Evaluators.

The combination space is large

A ledger with 200 skills, 6 model versions, 4 tool-version combinations and 20 task families has nearly 100,000 possible conditions. Few can be measured at useful sample sizes. Three strategies make the ledger practical:

  1. Trigger-driven measurement. Measure when something changes: a new model is a candidate for deployment, a tool interface changes, a skill is revised. The Model Change Assurance process is the natural trigger for model changes.
  2. Risk-weighted sampling. Measure skills with consequential effects, high usage or prior instability first, and sample others.
  3. Partial pooling. Hierarchical models can borrow strength across related skills and families, producing more stable estimates for sparsely measured cells. They also make it explicit when an estimate rests mostly on related cells rather than direct measurement. Such estimates should be labeled as model-based, not measured.

What counts as the same skill?

Transfer can only be measured for something that stays the same while the condition changes. For a skill, that means the contract, what it promises, which tools and permissions it uses and how success is verified, must be held fixed while its realization may be regenerated for the new model. The ledger therefore records two kinds of target measurement. Frozen transfer uses the realization exactly as it was written for the source condition. Adapted transfer uses a new realization produced for the target condition by a documented procedure and budget. Frozen transfer answers "what happens if we change models and do nothing?". Adapted transfer answers "how much effort does it take to restore the skill's value?". Both matter, and conflating them hides the cost of migration.

A reporting standard

Each ledger report should include, for every skill and condition reported: the exact skill, model, tool and verifier versions; whether transfer was frozen or adapted, and the adaptation budget if adapted; the number of tasks and repeated trials per arm; success with and without the skill, with intervals; reliability across trials; cost per verified outcome in both arms; the derived status and the pre-registered rule that produced it; and the share of conditions in the reported scope that remain untested. The last item keeps the ledger honest. A report that lists only validated skills says nothing about the much larger space nobody has checked.

From measurements to decisions

Figure 3 shows an illustrative ledger view across model versions. The entries are invented to illustrate the reporting format.

Matrix with five illustrative skills as rows (invoice reconciliation, bug triage, contract clause extraction, ticket routing, data migration check) and four model versions as columns (development model, upgrade one, upgrade two, different provider). Cells use symbols: validated, degraded, negative transfer, obsolete because baseline already succeeds, and untested. Ticket routing becomes obsolete on later models; contract clause extraction shows negative transfer on one upgrade.

Figure 3. An illustrative transfer matrix. What a ledger view might look like across model versions for five skills. Entries are invented to illustrate the reporting format, including negative transfer and obsolescence; they are not Ethen measurements. Evidence label: ILLUSTRATIVE — NOT MEASURED ETHEN DATA. Source: Invented entries for format illustration only.

We propose status rules of the following form, with thresholds fixed in advance for each risk class:

  • Validated: target gain's lower confidence bound above zero, and reliability not materially below the no-skill baseline.
  • Degraded: target gain positive but retention below a stated fraction, or reliability materially reduced.
  • Negative transfer: target gain's upper confidence bound below zero. The skill must be disabled for that condition until a new realization passes.
  • Obsolete: target gain indistinguishable from zero with a high baseline. Candidate for retirement.
  • Untested: no measurement. Production use in an untested condition should be a deliberate, recorded decision rather than a default.

These statuses feed back into deployment. A skill marked negative for a model version is not loaded when that model is in use. A skill marked obsolete across all current models is retired. The benchmark track that operationalizes these measurements on standard tasks is VerifiedWork Transfer. The experiment for the high-stakes case of a frontier-model upgrade is How to Measure Whether AI Skills Survive a Frontier-Model Upgrade.

Why negative transfer deserves special attention

Negative transfer is the outcome most likely to go unnoticed. When a skill compensates for a weakness of an older model, by spelling out steps the older model skipped, constraining output formats it got wrong, or forbidding approaches it handled poorly, a newer model without that weakness may perform better without the skill. The skill now constrains a capable model to an older model's workaround. Aggregate success may still look acceptable, so nobody investigates. Only a paired comparison with the skill switched off reveals the harm. Recovery procedures are a special case of the same question; see Testing Whether Recovery Knowledge Transfers Across Tools.

Threats to validity

  • Replay or environment drift between measurements can change results independently of the skill. Snapshots and environment versions must be pinned.
  • Hidden model changes behind a stable version name can contaminate before-and-after comparisons. Where providers allow it, pin exact versions; where they do not, re-measure baselines at the same time as treatments.
  • Task-sample shift between source and target measurements confounds transfer with difficulty. Paired designs on the same tasks avoid it.
  • Skill-author bias. Authors may tune realizations on the evaluation tasks. Decisive measurements should use sealed task sets managed by someone other than the skill's author.

Limitations

This paper proposes a measurement framework. No ledger has been built and no transfer has been measured. The combination space means that most conditions will remain untested at any time, and model-based estimates for them carry real uncertainty. Thresholds for status decisions are judgments that must be set per risk class. The framework measures transfer of skills; transfer of fine-tuned weights or learned policies raises additional issues not treated here.

Conclusion

Every organization that builds on foundation models is betting that some of what it builds will outlast the next model. The Capability Transfer Ledger turns that bet into a record: for each skill and each change, did it keep helping, help less, stop mattering, or start hurting? Measured with and without the skill, on paired tasks, it is the evidence that distinguishes an accumulating asset from a recurring cost.

FAQ

What is negative transfer? A skill that makes performance worse in a new condition than having no skill at all, typically because it encodes a workaround the new model does not need.

Why measure without the skill? Because a new model may succeed on its own. Without the no-skill baseline, apparent transfer may be the model's improvement, not the skill's.

Does the ledger require Skill IR? No. Any skill that can be switched on and off on the same task can be measured.

References

  1. Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
  2. Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
  3. Zheng, B. et al. (2025). SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills. arXiv:2504.07079. https://arxiv.org/abs/2504.07079
  4. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  5. Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
  6. Yan, S. et al. (2020). Positive-Congruent Training: Towards Regression-Free Model Updates. arXiv:2011.09161. https://arxiv.org/abs/2011.09161
  7. Srivastava, M. et al. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
  8. Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
  9. Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
  10. McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2):153–157. https://doi.org/10.1007/BF02295996
  11. Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Models & Intelligence

    Why Ethen Sometimes Won’t Use the Newest Model

    Should you upgrade to the newest AI model? Not automatically. Ethen sometimes holds back from the newest model because a model that is better on average can still be worse on the specific work people rely on — and averages hide exactly those regressions. A new version may follow formats differently, refuse different requests, call tools differently, cost more per finished task, run slower, or come with different terms. Prompts and skills tuned for the previous model may perform worse until they are adapted. So Ethen treats a new model as a candidate, not an upgrade: we test it on representative work task by task, look past the average to regressions, cost and behavior, switch kind of work by kind of work where it wins, keep tested versions pinned and a fallback available, and keep watching after a switch. Sometimes that means adopting a new model within days. Sometimes it means not adopting it at all.

  • Models & Intelligence

    Why Local AI Still Matters in a Cloud-First World

    People run AI locally for five main reasons, and they hold even as cloud models get larger and cheaper. Work processed entirely on your own machine is not sent to a model provider. Local models keep working without a network connection. You decide when a local model changes, so its behavior does not shift under you. Costs are paid up front in hardware rather than growing with every request. And local runtimes are an open playground for experimenting with models. The trade-offs are real: local models are smaller than the largest hosted ones, quality depends on your hardware, you maintain the setup yourself, and the privacy benefit applies only to the steps that actually stay on your machine. This article explains when local AI is worth it, what it does and does not protect, and how Ethen approaches it.

  • Company

    What Ethen Research Lab Is Exploring Beyond AI Products

    Ethen Research Lab research areas reach beyond any single product. The Lab is organized into eight published programs — Evaluation and Verification; Trust and Accountable AI Work; Adaptive Intelligence; Context, Skills and Transfer; Model Intelligence and Faros; Data and Learning Systems; Enterprise and Sovereign AI; and AI Security — connected by one thread: how AI systems can learn from work that can be checked. Some questions feed directly into products. Others look further out: how to measure whether automated graders can be trusted, whether skills survive when the underlying model changes, whether organizations can improve AI without exporting their data, whether agents can predict the effects of their actions, and whether automated research can keep hypotheses separate from confirmed findings. Almost all of this work is published as proposals, protocols and benchmark designs; very little is results. This article describes the questions honestly, without claims the evidence cannot support.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.