Research Protocol · 2026-10-03 · Context / Skills / Transfer
How to Measure Whether AI Skills Survive a Frontier-Model Upgrade
Every frontier release is a natural experiment on the skills built for the model before it. This protocol describes how to run that experiment deliberately, rather than discovering the results in production.
Abstract
Organizations encode much of their agent know-how as skills: procedures, instructions, examples, tool wrappers and verifiers tuned for particular work on a particular model. Whether that investment survives the next frontier-model upgrade is usually assumed, not measured. AI skill portability is the property this protocol is designed to measure. It treats each frontier release as a pre-registered event study. Before the release, a frozen battery of skills is measured on the incumbent model with and without each skill, across task families and tool versions. After the release, the same battery is run on the new model, first frozen and then after a fixed adaptation budget, and again some weeks later to detect drift. Each skill-by-condition cell is classified as transferred, degraded, negatively transferred, obsolete, or correctly abstained, and the results are assembled into a compatibility matrix with the cost of revalidation recorded alongside. The protocol builds on the Skill IR representation, the estimands of the Capability Transfer Ledger and the public design of VerifiedWork Transfer. No measurements have been made. A working title that implied results was replaced because none exist.
Why an upgrade is the right moment to measure
Skills are usually built against one model and tested against that model. When a stronger model arrives, three outcomes are plausible and all are observed in adjacent literatures. The skill may still help. It may stop mattering, because the new model can do the work unaided. Or it may hurt: instructions written to compensate for an old model's weaknesses can constrain a model that no longer has them. The backward-compatibility literature documents instance-level regressions when models are updated even as average accuracy rises (Srivastava et al.; Echterhoff et al.), and prompt formatting choices that look immaterial can change measured performance substantially (Sclar et al.).
There is also positive evidence that skills can transfer between agents. SkillWeaver reports that skills synthesized as APIs by stronger web agents substantially improved weaker agents on WebArena (Zheng et al.). Skill libraries in embodied settings (Wang et al., 2023) and induced workflow memories for web agents (Wang et al., 2024) point in the same direction. None of these settles the question for a specific organization's skills when its incumbent model is replaced by a stronger one. That question is an empirical one, and an upgrade is the moment when it can be asked cleanly: the old model is still available, the new one has just arrived, and nothing has yet been adapted.
What is being measured
The object of study is the outcome of a function:
skill identity × model version × tool version × task family → transfer result
Skill identity. A skill is identified by the content hash of its contract and realization, in the sense of Skill IR: its declared inputs, outputs, preconditions, tools, permissions, verifier and applicability boundary, and the model-specific instructions or code that implement it. Two skills with the same name but different realizations are different skills.
Model version. The incumbent and the new model, recorded with all identifiers and version metadata the provider exposes. Where a provider offers dated snapshots, the snapshot is pinned.
Tool version. The interface versions of the tools each skill depends on. The protocol includes at least one deliberate tool-version change per tool family, because a model upgrade and an API revision often arrive in the same quarter and their effects must not be confused.
Task family. Sets of tasks that share tools, verification and risk profile. Each skill is tested on tasks inside its declared boundary and on a smaller set just outside it.
Design: a pre-registered event study
Figure 1 shows the conditions and the timeline.
Figure 1. A pre-registered event study around one upgrade. Each skill is measured with and without itself on both models. Source gain comes from the incumbent; target gain from the new model, first frozen and then after a fixed adaptation budget. A later anchor re-run detects drift behind a stable model name. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Pre-release window. The skill battery, task sets, verifiers, hypotheses and classification thresholds are frozen and registered before the new model is available. Each skill is measured on the incumbent with and without the skill. This produces the source gain: how much the skill helps the model it was built for.
Release, frozen. As soon as the new model is available, the battery is run on it with and without each skill, unchanged. This produces the target gain. Comparing the two is the central measurement.
Release, adapted. Each skill is then given a fixed adaptation budget, stated in engineer-hours, automated optimization calls and evaluation runs, and re-measured. The difference between frozen and adapted results, together with the budget actually spent, is the revalidation cost.
Release plus several weeks. A fixed anchor subset is re-run to detect behavior changes behind a stable model identifier (Chen et al.). If anchor results move beyond their noise band, later results are reported separately.
Within each cell, every task is run several times so that reliability across trials (pass^k; Yao et al.) can be estimated and stochastic flips separated from systematic change.
Classifying each cell
Figure 2 defines the outcome classes.
Figure 2. Classifying each skill after the upgrade. Classes are defined by the source gain on the incumbent and the target gain on the new model, each with a paired interval. Negative transfer and boundary failures are visible only because the no-skill condition is run in every cell. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Let Gs be the source gain (incumbent with skill minus incumbent without) and Gt the target gain (new model with skill minus new model without), each with a paired confidence interval.
- Transferred. Gt is positive and its lower bound is at least a pre-registered fraction of Gs. The skill still adds value of a similar order.
- Degraded. Gt is positive but clearly smaller than Gs.
- Obsolete. Gt is indistinguishable from zero and the new model without the skill matches or exceeds the incumbent with it. The skill no longer earns its maintenance cost.
- Negative transfer. Gt is negative with an interval excluding zero. The skill makes the new model worse than it would be alone. This is the costliest outcome, because it is invisible unless the no-skill condition is run.
- Abstention. On tasks outside the declared boundary, the skill should decline to apply. Correct abstention is a success; applying the skill where it should have abstained is scored as a boundary failure.
A "transferred" classification requires an equivalence-style test against the pre-registered margin (Schuirmann), not merely the absence of a significant difference. Cells without enough tasks to classify are reported as inconclusive.
The same machinery measures safety. A skill that raises verified success on the new model while increasing attempted out-of-scope actions, checked deterministically against the governing mandate, is reported as negative transfer on the safety axis regardless of its success gain.
Applicability boundaries
Skills are not meant to apply everywhere, and much of their risk lies at the edges. Each skill under test declares its applicability boundary: the task families, input conditions and tool versions for which it claims to work. The protocol tests the boundary in both directions. Inside the boundary, it measures gain. Just outside it, with tasks that are similar but violate a precondition, it measures whether the skill or the agent using it abstains. Upgrades can move boundaries. A stronger model may handle cases the old skill excluded, so a boundary that was correct for the incumbent becomes needlessly narrow. Or a new model may apply a skill more eagerly, so that out-of-boundary use rises. Both are reported.
Separating capability from compatibility
A negative or degraded result has several possible causes, and the design includes controls for each.
Prompt-format sensitivity. A skill may fail on the new model because of formatting, not substance. Each skill is run under a small set of semantically equivalent format variants, and results are reported as a range. A skill whose result flips across equivalent formats is reported as format-sensitive rather than as transferred or failed (Sclar et al.).
Tool-version confounding. Because tool versions are crossed with model versions in a subset of cells, the effect of an API change can be estimated separately from the effect of the model change.
Verifier change. Verifier versions are pinned for the study. If a verifier is itself model-based and the new model is also the judge's model family, a second judge from a different family is used for a sample, since judges can favor their own generations (Panickssery et al.). Verifier and task-setup flaws can misstate agent performance substantially even on established benchmarks (Zhu et al.), so each family's verifier is audited on a sample before the study begins.
Ceiling effects. If the new model without skills is near the verifier's ceiling on a family, gains cannot be measured there. Such families are reported as saturated, not as obsolescence.
Contamination. If a skill's content or task set was public before the new model's training cutoff, its apparent transfer may reflect memorization. Contamination of public test sets has been shown to inflate measured performance (Zhang et al.), so private, post-cutoff task sets are used for confirmatory cells.
The compatibility matrix
Figure 3 shows the output format.
Figure 3. The compatibility matrix (illustrative layout). Rows are skills; columns are model and tool versions. Each cell carries a class, an interval, task and trial counts, and the adaptation budget spent. Symbols here show the layout only; no cell reports a measurement. Evidence label: ILLUSTRATIVE LAYOUT — NO DATA. Source: Ethen research protocol (proposed); illustrative layout, no data.
The deliverable is a matrix with skills as rows and model-and-tool versions as columns, each cell holding one of the classes above, its interval, the number of tasks and trials, and the adaptation budget spent. The matrix is accompanied by a short report per skill that explains every negative-transfer or boundary-failure cell. This matrix is the entry format for the Capability Transfer Ledger, which accumulates such measurements across releases.
From the matrix, a skill owner can make three decisions: keep the skill unchanged for the new model, revise it within the measured adaptation cost, or retire it. Retirement is a legitimate and often desirable outcome. A skill that the new model has made obsolete is maintenance cost and attack surface with no benefit.
Sampling a large space
A full cross of skills, models, tool versions, families and format variants is too large to run exhaustively. The protocol prioritizes in three ways. Skills are ordered by usage and by the risk of the work they touch, and the highest-priority skills receive full crossing. Lower-priority skills receive the frozen with-and-without comparison on their primary family only. Tool-version crossing is limited to the tools most often changed. The selection rule is registered in advance, so that the matrix does not silently concentrate on skills expected to do well.
Statistical analysis
Within each cell, paired task-level outcomes are compared with methods for paired binary data, and intervals are computed by cluster bootstrap over tasks grouped by source. The key quantity, the change from Gs to Gt, is an interaction contrast, which typically needs more data than either gain alone; power planning uses pilot estimates of discordance and between-trial variance. Classification across many cells is controlled for false discovery (Benjamini & Hochberg). Revalidation cost is reported descriptively, with the adaptation budget and the gain recovered.
Relation to adjacent work in this series
This protocol sits between three others. A Research Protocol for Model Change Assurance asks whether a whole configuration change will regress an organization's work; this protocol asks which individual skills survive the change. VerifiedWork Transfer asks the same question on shared public tasks, so that methods can be compared across organizations. The Capability Transfer Ledger stores the results over time. Reporting cost alongside success, a recommendation of recent work on agent evaluation (Kapoor et al.), applies here as well: a skill that transfers only at high adaptation cost may still not be worth keeping.
Limitations
The protocol has not been run, and its classification margins are proposed. Its results apply to the skills, models, tool versions and task families tested, and a favorable result for one upgrade does not predict the next. Frontier releases arrive on the provider's schedule, which limits pre-registration time. Adaptation budgets are hard to standardize across teams. Measuring negative transfer reliably requires the no-skill condition in every cell, which roughly doubles the cost of the study, and that cost is deliberate.
Conclusion
Skills are an investment made against one model. Each frontier upgrade either preserves, erodes, eliminates or reverses that investment, and only measurement with and without the skill can tell which. This protocol turns each release into a pre-registered study whose output is a compatibility matrix and a revalidation cost, so that skill owners can keep, revise or retire their work on evidence rather than habit.
FAQ
Why run every skill with and without the skill on the new model? Without the no-skill condition, negative transfer and obsolescence are invisible. A skill can look fine while making the new model worse than it would be alone.
Is retiring a skill a failure? No. If the new model no longer needs it, retirement removes maintenance cost and attack surface.
How is this different from the Capability Transfer Ledger? The ledger is the method and the record. This protocol is the study design for one upgrade event, whose results enter the ledger.
Related research
- Skill IR: Toward Model-Independent Agent Capabilities — Skill IR.
- The Capability Transfer Ledger: Measuring Whether AI Skills Survive Model Upgrades — ledger.
- VerifiedWork Transfer: Measuring Capability Across Model and Tool Changes — benchmark track.
- A Research Protocol for Model Change Assurance — MCA protocol.
References
- Srivastava, M. et al. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
- Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
- Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
- Zheng, B. et al. (2025). SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills. arXiv:2504.07079. https://arxiv.org/abs/2504.07079
- Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
- Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
- Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
- Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. J. Pharmacokinet. Biopharm. 15:657–680. https://doi.org/10.1007/BF01068419
- Panickssery, A. et al. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076. https://arxiv.org/abs/2404.13076
- Zhu, Y. et al. (2025). Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825. https://arxiv.org/abs/2507.02825
- Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
- Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
- Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Skill IR: Toward Model-Independent Agent Capabilities
A research proposal for Skill IR: AI agent skills as portable, testable capability contracts with permissions, verifiers and compatibility records.
- The Capability Transfer Ledger: Measuring Whether AI Skills Survive Model Upgrades
A methods paper on measuring capability transfer: whether AI agent skills keep working across model, tool and task changes, including negative transfer.
- Evidence-Preserving Context: Compressing Agent Memory Without Losing Obligations
A research proposal for context compaction that keeps obligations, permissions, deadlines and evidence out of lossy summaries while cutting an agent's token cost.
Explained on the Ethen Blog
- Why Ethen Sometimes Won’t Use the Newest Model
Should you upgrade to the newest AI model? Not automatically. Ethen sometimes holds back from the newest model because a model that is better on average can still be worse on the specific work people rely on — and averages hide exactly those regressions. A new version may follow formats differently, refuse different requests, call tools differently, cost more per finished task, run slower, or come with different terms. Prompts and skills tuned for the previous model may perform worse until they are adapted. So Ethen treats a new model as a candidate, not an upgrade: we test it on representative work task by task, look past the average to regressions, cost and behavior, switch kind of work by kind of work where it wins, keep tested versions pinned and a fallback available, and keep watching after a switch. Sometimes that means adopting a new model within days. Sometimes it means not adopting it at all.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.