Skip to content

EthenEthenEthen

Research Protocol · 2026-10-03 · Context / Skills / Transfer

How to Measure Whether AI Skills Survive a Frontier-Model Upgrade

Publication type
Research Protocol
Evidence status
Protocol / Planned Experiment: The method is defined. The experiment has not been run, so no results are reported.research protocol; not yet run; no skill-survival measurements exist
Research program
Context / Skills / Transfer
Published
Authors
Ethen Research Lab
Reading time
12 min read

Every frontier release is a natural experiment on the skills built for the model before it. This protocol describes how to run that experiment deliberately, rather than discovering the results in production.

Cover image for "How to Measure Whether AI Skills Survive a Frontier-Model Upgrade". Decorative abstract motif; contains no data.

Abstract

Organizations encode much of their agent know-how as skills: procedures, instructions, examples, tool wrappers and verifiers tuned for particular work on a particular model. Whether that investment survives the next frontier-model upgrade is usually assumed, not measured. AI skill portability is the property this protocol is designed to measure. It treats each frontier release as a pre-registered event study. Before the release, a frozen battery of skills is measured on the incumbent model with and without each skill, across task families and tool versions. After the release, the same battery is run on the new model, first frozen and then after a fixed adaptation budget, and again some weeks later to detect drift. Each skill-by-condition cell is classified as transferred, degraded, negatively transferred, obsolete, or correctly abstained, and the results are assembled into a compatibility matrix with the cost of revalidation recorded alongside. The protocol builds on the Skill IR representation, the estimands of the Capability Transfer Ledger and the public design of VerifiedWork Transfer. No measurements have been made. A working title that implied results was replaced because none exist.

Why an upgrade is the right moment to measure

Skills are usually built against one model and tested against that model. When a stronger model arrives, three outcomes are plausible and all are observed in adjacent literatures. The skill may still help. It may stop mattering, because the new model can do the work unaided. Or it may hurt: instructions written to compensate for an old model's weaknesses can constrain a model that no longer has them. The backward-compatibility literature documents instance-level regressions when models are updated even as average accuracy rises (Srivastava et al.; Echterhoff et al.), and prompt formatting choices that look immaterial can change measured performance substantially (Sclar et al.).

There is also positive evidence that skills can transfer between agents. SkillWeaver reports that skills synthesized as APIs by stronger web agents substantially improved weaker agents on WebArena (Zheng et al.). Skill libraries in embodied settings (Wang et al., 2023) and induced workflow memories for web agents (Wang et al., 2024) point in the same direction. None of these settles the question for a specific organization's skills when its incumbent model is replaced by a stronger one. That question is an empirical one, and an upgrade is the moment when it can be asked cleanly: the old model is still available, the new one has just arrived, and nothing has yet been adapted.

What is being measured

The object of study is the outcome of a function:

skill identity × model version × tool version × task family → transfer result

Skill identity. A skill is identified by the content hash of its contract and realization, in the sense of Skill IR: its declared inputs, outputs, preconditions, tools, permissions, verifier and applicability boundary, and the model-specific instructions or code that implement it. Two skills with the same name but different realizations are different skills.

Model version. The incumbent and the new model, recorded with all identifiers and version metadata the provider exposes. Where a provider offers dated snapshots, the snapshot is pinned.

Tool version. The interface versions of the tools each skill depends on. The protocol includes at least one deliberate tool-version change per tool family, because a model upgrade and an API revision often arrive in the same quarter and their effects must not be confused.

Task family. Sets of tasks that share tools, verification and risk profile. Each skill is tested on tasks inside its declared boundary and on a smaller set just outside it.

Design: a pre-registered event study

Figure 1 shows the conditions and the timeline.

Timeline of four stages. Pre-release: battery, tasks, verifiers and thresholds frozen; incumbent measured with and without each skill, giving source gain. Release frozen: new model measured with and without each unchanged skill, giving target gain. Release adapted: each skill revised within a fixed adaptation budget and re-measured, giving revalidation cost. Release plus several weeks: anchor subset re-run to detect drift. Below, a two-by-two of conditions per cell: incumbent without skill, incumbent with skill, new model without skill, new model with skill.

Figure 1. A pre-registered event study around one upgrade. Each skill is measured with and without itself on both models. Source gain comes from the incumbent; target gain from the new model, first frozen and then after a fixed adaptation budget. A later anchor re-run detects drift behind a stable model name. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Pre-release window. The skill battery, task sets, verifiers, hypotheses and classification thresholds are frozen and registered before the new model is available. Each skill is measured on the incumbent with and without the skill. This produces the source gain: how much the skill helps the model it was built for.

Release, frozen. As soon as the new model is available, the battery is run on it with and without each skill, unchanged. This produces the target gain. Comparing the two is the central measurement.

Release, adapted. Each skill is then given a fixed adaptation budget, stated in engineer-hours, automated optimization calls and evaluation runs, and re-measured. The difference between frozen and adapted results, together with the budget actually spent, is the revalidation cost.

Release plus several weeks. A fixed anchor subset is re-run to detect behavior changes behind a stable model identifier (Chen et al.). If anchor results move beyond their noise band, later results are reported separately.

Within each cell, every task is run several times so that reliability across trials (pass^k; Yao et al.) can be estimated and stochastic flips separated from systematic change.

Classifying each cell

Figure 2 defines the outcome classes.

Table of six outcome classes with their definition and the usual decision. Transferred: target gain positive and at least a pre-registered fraction of source gain; keep. Degraded: target gain positive but clearly smaller; revise or keep if cheap. Obsolete: target gain indistinguishable from zero while the new model alone matches the incumbent with the skill; retire. Negative transfer: target gain negative with interval excluding zero; retire or rewrite, highlighted. Correct abstention: skill declines outside its declared boundary; keep. Boundary failure: skill applied outside its boundary; fix the boundary.

Figure 2. Classifying each skill after the upgrade. Classes are defined by the source gain on the incumbent and the target gain on the new model, each with a paired interval. Negative transfer and boundary failures are visible only because the no-skill condition is run in every cell. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Let Gs be the source gain (incumbent with skill minus incumbent without) and Gt the target gain (new model with skill minus new model without), each with a paired confidence interval.

  • Transferred. Gt is positive and its lower bound is at least a pre-registered fraction of Gs. The skill still adds value of a similar order.
  • Degraded. Gt is positive but clearly smaller than Gs.
  • Obsolete. Gt is indistinguishable from zero and the new model without the skill matches or exceeds the incumbent with it. The skill no longer earns its maintenance cost.
  • Negative transfer. Gt is negative with an interval excluding zero. The skill makes the new model worse than it would be alone. This is the costliest outcome, because it is invisible unless the no-skill condition is run.
  • Abstention. On tasks outside the declared boundary, the skill should decline to apply. Correct abstention is a success; applying the skill where it should have abstained is scored as a boundary failure.

A "transferred" classification requires an equivalence-style test against the pre-registered margin (Schuirmann), not merely the absence of a significant difference. Cells without enough tasks to classify are reported as inconclusive.

The same machinery measures safety. A skill that raises verified success on the new model while increasing attempted out-of-scope actions, checked deterministically against the governing mandate, is reported as negative transfer on the safety axis regardless of its success gain.

Applicability boundaries

Skills are not meant to apply everywhere, and much of their risk lies at the edges. Each skill under test declares its applicability boundary: the task families, input conditions and tool versions for which it claims to work. The protocol tests the boundary in both directions. Inside the boundary, it measures gain. Just outside it, with tasks that are similar but violate a precondition, it measures whether the skill or the agent using it abstains. Upgrades can move boundaries. A stronger model may handle cases the old skill excluded, so a boundary that was correct for the incumbent becomes needlessly narrow. Or a new model may apply a skill more eagerly, so that out-of-boundary use rises. Both are reported.

Separating capability from compatibility

A negative or degraded result has several possible causes, and the design includes controls for each.

Prompt-format sensitivity. A skill may fail on the new model because of formatting, not substance. Each skill is run under a small set of semantically equivalent format variants, and results are reported as a range. A skill whose result flips across equivalent formats is reported as format-sensitive rather than as transferred or failed (Sclar et al.).

Tool-version confounding. Because tool versions are crossed with model versions in a subset of cells, the effect of an API change can be estimated separately from the effect of the model change.

Verifier change. Verifier versions are pinned for the study. If a verifier is itself model-based and the new model is also the judge's model family, a second judge from a different family is used for a sample, since judges can favor their own generations (Panickssery et al.). Verifier and task-setup flaws can misstate agent performance substantially even on established benchmarks (Zhu et al.), so each family's verifier is audited on a sample before the study begins.

Ceiling effects. If the new model without skills is near the verifier's ceiling on a family, gains cannot be measured there. Such families are reported as saturated, not as obsolescence.

Contamination. If a skill's content or task set was public before the new model's training cutoff, its apparent transfer may reflect memorization. Contamination of public test sets has been shown to inflate measured performance (Zhang et al.), so private, post-cutoff task sets are used for confirmatory cells.

The compatibility matrix

Figure 3 shows the output format.

Illustrative matrix with five hypothetical skills as rows (invoice reconciliation, bug triage, contract amendment draft, CRM record update, research synthesis) and four columns (incumbent model with tool v1; new model frozen with tool v1; new model adapted with tool v1; new model adapted with tool v2). Cells contain placeholder symbols for transferred, degraded, obsolete, negative transfer and inconclusive. A legend explains the symbols and states that the layout contains no data.

Figure 3. The compatibility matrix (illustrative layout). Rows are skills; columns are model and tool versions. Each cell carries a class, an interval, task and trial counts, and the adaptation budget spent. Symbols here show the layout only; no cell reports a measurement. Evidence label: ILLUSTRATIVE LAYOUT — NO DATA. Source: Ethen research protocol (proposed); illustrative layout, no data.

The deliverable is a matrix with skills as rows and model-and-tool versions as columns, each cell holding one of the classes above, its interval, the number of tasks and trials, and the adaptation budget spent. The matrix is accompanied by a short report per skill that explains every negative-transfer or boundary-failure cell. This matrix is the entry format for the Capability Transfer Ledger, which accumulates such measurements across releases.

From the matrix, a skill owner can make three decisions: keep the skill unchanged for the new model, revise it within the measured adaptation cost, or retire it. Retirement is a legitimate and often desirable outcome. A skill that the new model has made obsolete is maintenance cost and attack surface with no benefit.

Sampling a large space

A full cross of skills, models, tool versions, families and format variants is too large to run exhaustively. The protocol prioritizes in three ways. Skills are ordered by usage and by the risk of the work they touch, and the highest-priority skills receive full crossing. Lower-priority skills receive the frozen with-and-without comparison on their primary family only. Tool-version crossing is limited to the tools most often changed. The selection rule is registered in advance, so that the matrix does not silently concentrate on skills expected to do well.

Statistical analysis

Within each cell, paired task-level outcomes are compared with methods for paired binary data, and intervals are computed by cluster bootstrap over tasks grouped by source. The key quantity, the change from Gs to Gt, is an interaction contrast, which typically needs more data than either gain alone; power planning uses pilot estimates of discordance and between-trial variance. Classification across many cells is controlled for false discovery (Benjamini & Hochberg). Revalidation cost is reported descriptively, with the adaptation budget and the gain recovered.

Relation to adjacent work in this series

This protocol sits between three others. A Research Protocol for Model Change Assurance asks whether a whole configuration change will regress an organization's work; this protocol asks which individual skills survive the change. VerifiedWork Transfer asks the same question on shared public tasks, so that methods can be compared across organizations. The Capability Transfer Ledger stores the results over time. Reporting cost alongside success, a recommendation of recent work on agent evaluation (Kapoor et al.), applies here as well: a skill that transfers only at high adaptation cost may still not be worth keeping.

Limitations

The protocol has not been run, and its classification margins are proposed. Its results apply to the skills, models, tool versions and task families tested, and a favorable result for one upgrade does not predict the next. Frontier releases arrive on the provider's schedule, which limits pre-registration time. Adaptation budgets are hard to standardize across teams. Measuring negative transfer reliably requires the no-skill condition in every cell, which roughly doubles the cost of the study, and that cost is deliberate.

Conclusion

Skills are an investment made against one model. Each frontier upgrade either preserves, erodes, eliminates or reverses that investment, and only measurement with and without the skill can tell which. This protocol turns each release into a pre-registered study whose output is a compatibility matrix and a revalidation cost, so that skill owners can keep, revise or retire their work on evidence rather than habit.

FAQ

Why run every skill with and without the skill on the new model? Without the no-skill condition, negative transfer and obsolescence are invisible. A skill can look fine while making the new model worse than it would be alone.

Is retiring a skill a failure? No. If the new model no longer needs it, retirement removes maintenance cost and attack surface.

How is this different from the Capability Transfer Ledger? The ledger is the method and the record. This protocol is the study design for one upgrade event, whose results enter the ledger.

References

  1. Srivastava, M. et al. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
  2. Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
  3. Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
  4. Zheng, B. et al. (2025). SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills. arXiv:2504.07079. https://arxiv.org/abs/2504.07079
  5. Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
  6. Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
  7. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  8. Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
  9. Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. J. Pharmacokinet. Biopharm. 15:657–680. https://doi.org/10.1007/BF01068419
  10. Panickssery, A. et al. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076. https://arxiv.org/abs/2404.13076
  11. Zhu, Y. et al. (2025). Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825. https://arxiv.org/abs/2507.02825
  12. Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
  13. Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
  14. Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Models & Intelligence

    Why Ethen Sometimes Won’t Use the Newest Model

    Should you upgrade to the newest AI model? Not automatically. Ethen sometimes holds back from the newest model because a model that is better on average can still be worse on the specific work people rely on — and averages hide exactly those regressions. A new version may follow formats differently, refuse different requests, call tools differently, cost more per finished task, run slower, or come with different terms. Prompts and skills tuned for the previous model may perform worse until they are adapted. So Ethen treats a new model as a candidate, not an upgrade: we test it on representative work task by task, look past the average to regressions, cost and behavior, switch kind of work by kind of work where it wins, keep tested versions pinned and a fallback available, and keep watching after a switch. Sometimes that means adopting a new model within days. Sometimes it means not adopting it at all.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.