Skip to content

EthenEthenEthen

Benchmark Design · 2026-10-03 · Evaluation & Verification

VerifiedWork Transfer: Measuring Capability Across Model and Tool Changes

Publication type
Benchmark Design
Research program
Evaluation & Verification
Published
Authors
Ethen Research Lab
Reading time
12 min read

An agent system is built against one model, one set of tool interfaces and one mix of tasks. All three will change. This benchmark track measures what survives.

Cover image for "VerifiedWork Transfer: Measuring Capability Across Model and Tool Changes". Decorative abstract motif; contains no data.

Abstract

Agent systems accumulate skills, prompts, configurations and recovery procedures tuned to the conditions in which they were developed: a particular model version, provider, set of tool interfaces, task mix and runtime. In deployment, every one of those conditions changes. Models are upgraded, providers switched, APIs revised, new kinds of work arrive. VerifiedWork Transfer is the benchmark track of Ethen VerifiedWork that measures what survives those changes. Submissions are developed in public source conditions and evaluated in hidden target conditions that differ along five axes: model version, model provider, tool-interface version, task family and harness version. Each submission is evaluated twice, frozen as developed and adapted after a fixed, reported adaptation budget. Results are reported as a transfer matrix with gain, retention, negative transfer, obsolescence and adaptation cost, not as a single score. Tool-interface drift is generated deliberately, including drift that changes semantics, and silent misuse of changed semantics is scored as a critical failure. This capability transfer benchmark is a proposed design that has not been run.

The question this track answers

Research on experience-driven agents shows that skills and workflows extracted from past work can improve later performance (Wang et al., Voyager; Wang et al., Agent Workflow Memory; Zheng et al., SkillWeaver). Deployed systems need to know something those studies rarely measure: whether that improvement survives when the conditions it was built in change. Hosted model behavior shifts between versions (Chen et al.). Prompt performance depends on formatting choices that carry no meaning (Sclar et al.). Model updates introduce instance-level regressions even as aggregate performance improves (Yan et al.; Echterhoff et al.).

The Capability Transfer Ledger, described in The Capability Transfer Ledger, is the methodology for measuring transfer of any skill in any condition. This track makes that methodology into a standardized, comparable benchmark: fixed task families, fixed source conditions, hidden target conditions and a common reporting format, so that different approaches to building transferable capability can be compared directly.

Task families and source conditions

The initial design draws on four task families chosen for strong verification: code maintenance (dependency updates, test repair, small migrations), graded by test suites and build checks; support operations (refunds, account changes, ticket resolution), graded by end-state checks on simulated systems; data reconciliation (matching invoices to orders, deduplicating records), graded by deterministic recomputation; and research evidence (answering questions from a frozen source collection with supported claims), graded by citation checks and calibrated rubrics. Each family has a public development split in the source condition and sealed splits for every target condition.

Source conditions are published in full: model versions, tool-interface versions, harness version and task families. Target conditions are fixed before submissions open, recorded with a cryptographic commitment, and revealed only after the release cycle, so that no one, including the benchmark operators, can adjust them in response to results.

Five axes of change

Target conditions differ from source conditions along five axes (Figure 1).

Source condition box on the left. Five arrows to five change axes in the middle: model version, model provider, tool interface version, task family, and harness or runtime version. Each leads to hidden target conditions on the right. A note says single-axis changes come first, then pairwise combinations.

Figure 1. Five axes of change between source and target conditions. A submission is developed in public source conditions and evaluated in hidden target conditions that differ along one axis at a time, then along combinations. Varying one axis at a time is what makes failures attributable. Evidence label: EXPERIMENT DESIGN. Source: Ethen benchmark design (proposed).

Model version. A newer or older version of the same provider's model. This is the most common change in practice.

Model provider. A model from a different provider, with different prompting conventions, tool-call formats and strengths.

Tool-interface version. The same tools with revised interfaces: renamed fields, new parameters, changed defaults or changed semantics. This axis is generated deliberately (see below).

Task family. A family of tasks not seen in development, sharing some structure with development families. This tests whether skills generalize beyond the tasks they were built on.

Harness or runtime version. A change in the surrounding execution system, such as context assembly, tool exposure or retry policy, holding the model fixed.

Each axis is first varied alone, so that failures can be attributed, and then in pairwise combinations, because real changes often arrive together.

Generating tool drift

Tool-interface drift is the least studied axis and among the most common in practice. The track generates it deliberately from versioned tool definitions, such as interface schemas published through the Model Context Protocol (Figure 2).

Table of six drift types: renamed field, new required parameter, changed pagination, changed default value, changed unit or meaning of a field, removed endpoint with replacement. Columns: semantics preserved or changed, and expected behavior. For a changed unit or meaning, expected behavior is to detect and adapt or stop; proceeding silently is a critical failure.

Figure 2. Tool-interface drift types and expected behavior. Tool drift is generated deliberately. Semantics-preserving drift should be absorbed; semantics-changing drift should be detected, with the agent adapting correctly or stopping. Proceeding silently on changed semantics is the critical failure. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen benchmark design (proposed).

The key distinction is between drift that preserves semantics, such as a renamed field or changed pagination, and drift that changes semantics, such as a field whose unit changes from cents to dollars, or a default that flips. Preserved-semantics drift should be absorbed: a robust system adapts and continues correctly. Changed-semantics drift should be detected: a robust system notices the change and either adapts correctly or stops and reports. The critical failure is proceeding silently on changed semantics, such as issuing a refund in the wrong unit. That failure is scored separately and prominently.

Skills that declare their tool dependencies by interface version, as proposed in Skill IR, give a system a mechanism for detecting such changes. The track measures whether that mechanism, or any other, works in practice.

Submission protocol

Figure 3 shows the protocol.

Flow: develop on public source conditions; submit skills and configuration; frozen evaluation in hidden targets; adaptation phase with a fixed, reported budget of compute and human time; adapted evaluation in the same targets; transfer report with target gain, retention, negative transfer, obsolescence and adaptation cost for each condition.

Figure 3. Submission protocol: frozen, then adapted. Each submission is evaluated twice in hidden target conditions: frozen, exactly as developed, and adapted, after a fixed and reported adaptation budget. Comparing the two separates robustness from the cost of migration. Evidence label: EXPERIMENT DESIGN. Source: Ethen benchmark design (proposed).

  1. Develop using public source conditions: a specified model version, tool-interface version, task families and harness.
  2. Submit the skills, prompts, configurations and any other artifacts the system uses.
  3. Frozen evaluation. The benchmark runs the submission, unchanged, in hidden target conditions.
  4. Adaptation phase. The submitter, or an automated procedure, may adapt the submission to each target condition within a fixed budget of compute and human time, which is reported. The adaptation phase sees only a development sample of target tasks, never the sealed evaluation tasks.
  5. Adapted evaluation on the sealed target tasks.
  6. Transfer report.

Frozen and adapted results answer different questions. Frozen transfer answers "what breaks if nobody does anything?". Adapted transfer answers "how much does it cost to restore performance?". An approach with weaker frozen transfer but very cheap adaptation may be preferable in practice to one with stronger frozen transfer and expensive adaptation.

Metrics

For each submission and condition, measured on paired tasks with and without the submission's artifacts, the report gives:

  • Target gain: verified success with the submission minus verified success without it, in the target condition. The no-artifact baseline is essential; without it, a better model's own improvement is mistaken for transfer.
  • Retention: target gain as a share of source gain, reported only where source gain is reliably positive.
  • Negative transfer: conditions where target gain is reliably below zero, meaning the submission makes the new condition worse than using no artifacts.
  • Obsolescence: conditions where target gain is near zero because the baseline already succeeds.
  • Critical drift failures: silent use of changed semantics.
  • Adaptation cost: compute and human time spent in the adaptation phase.
  • Reliability: pass^k across repeated trials (Yao et al.), since transfer can preserve average success while increasing variance.

Results are reported as a matrix of conditions, never collapsed into a single number. Different deployments care about different axes.

Baselines

Each release reports four reference submissions:

  • No artifacts: the model with tools and no skills, the baseline for every gain.
  • Prompt recipes: free-text instructions developed in the source condition.
  • Retrieved trajectories: examples of successful past work retrieved at run time, as in workflow-memory approaches.
  • Typed skills: contracts with declared preconditions, tool interfaces, permissions and verifiers, as in Skill IR.

The comparison of interest is whether richer representations transfer better, at what adaptation cost, and whether any of them avoid negative transfer.

Statistical design

Comparisons are paired on identical task instances and initial states. Intervals are computed with resampling at the task-family level to respect clustering. The track reports many conditions, so status decisions such as "negative transfer" control the false discovery rate across the conditions reported (Benjamini & Hochberg). Target conditions are hidden from submitters, but they are disclosed after each release cycle together with the full results, so the community can examine where transfer failed.

An illustrative transfer report

[ILLUSTRATIVE EXAMPLE — invented to show the format, not a result.] A submission built typed skills for support operations on a source model. In the frozen evaluation on a newer model from the same provider, its target gain is positive and retention high for refunds, but near zero for account changes, because the newer model handles them well without help: obsolescence, not failure. On a different provider's model, retention drops sharply for one skill whose instructions relied on a formatting convention the new model ignores. A modest adaptation budget restores most of it. Under a tool-interface change that switches an amount field from cents to dollars, the submission's refund skill detects the version mismatch through its declared interface dependency and stops. A prompt-recipe baseline proceeds silently and issues refunds at one hundredth of the intended value, a critical drift failure. The report shows each of these as a separate cell, with intervals, rather than averaging them into a score that would hide both the obsolescence and the critical failure.

Design choices considered and rejected

A single headline score. Rejected because it would average away the conditions that matter most to particular users, especially critical drift failures and negative transfer.

Random task splits. Rejected because random splits leak family structure into development and overstate generalization. Splits are by family and condition.

Revealing target conditions in advance. Rejected because submitters would, deliberately or not, tune to them. Targets are committed in advance and revealed afterward.

Frozen evaluation only. Rejected because it ignores the practical question of migration cost. Many approaches will be judged by how cheaply they can be adapted, not only by how robust they are untouched.

Measuring with the artifacts only. Rejected because without the no-artifact baseline in each target condition, model improvements masquerade as transfer and negative transfer is invisible.

Relation to the rest of VerifiedWork

This track uses the environments and reporting standard of Ethen VerifiedWork. Its frontier-model case is examined in depth in How to Measure Whether AI Skills Survive a Frontier-Model Upgrade, and its practical counterpart for a single organization is Model Change Assurance.

Threats to validity

Provider-side changes. A model served under a fixed version name can still change behind it. Where providers offer pinned snapshots, the track uses them; where they do not, baseline and treatment runs for a condition are interleaved in time so that drift affects both arms equally.

Adaptation leakage. If the adaptation phase can see sealed tasks, adapted results overstate transfer. Adaptation receives only a separate development sample from each target condition, and its access is logged.

Family definition. Whether a held-out family is "new" depends on how families were drawn. Families are defined before submissions open and documented, so that readers can judge how far each transfer reaches.

Grader drift. If graders change between source and target evaluations, they can create or hide apparent transfer. Grader versions are pinned per release and reported in every cell.

What this track cannot prove

A transfer result shows how a submission behaved in the hidden target conditions tested. It does not show how the submission will behave with model versions, providers or tool changes that did not exist when the benchmark was run. Future models may differ from today's in ways no target condition anticipated. Generated tool drift approximates real interface changes but cannot cover all of them.

Limitations

The track has not been built or run. Hidden target conditions require access to multiple model versions and providers, which depends on their availability and terms. Generated tool drift is a simplification of real API evolution. The adaptation budget is hard to standardize when adaptation involves human effort. Task-family transfer depends on how families are defined, which involves judgment.

Conclusion

Every capability an agent system accumulates is a bet that it will survive the next change. VerifiedWork Transfer measures that bet directly: develop against known conditions, evaluate against hidden ones, report what was retained, what was lost, what became harmful and what it cost to restore. The result is a map of where capability transfers, not a leaderboard.

FAQ

Why evaluate both frozen and adapted? Frozen results show what breaks without intervention. Adapted results show the cost of restoring performance. Both matter when deciding how to build capabilities.

What is a semantics-changing tool drift? A change that keeps an interface looking similar but changes what a value means, such as a field switching from cents to dollars. Proceeding silently after such a change is scored as a critical failure.

Why not report a single transfer score? Because deployments differ in which changes matter to them. A single number would hide the conditions where a submission fails or causes harm.

References

  1. Wang, G. et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291. https://arxiv.org/abs/2305.16291
  2. Wang, Z. Z. et al. (2024). Agent Workflow Memory. arXiv:2409.07429. https://arxiv.org/abs/2409.07429
  3. Zheng, B. et al. (2025). SkillWeaver. arXiv:2504.07079. https://arxiv.org/abs/2504.07079
  4. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  5. Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
  6. Yan, S. et al. (2020). Positive-Congruent Training: Towards Regression-Free Model Updates. arXiv:2011.09161. https://arxiv.org/abs/2011.09161
  7. Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
  8. Yao, S. et al. (2024). τ-bench. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
  9. Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
  10. Model Context Protocol. Specification. https://modelcontextprotocol.io/specification

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Company

    From Screen to Physical World: How We Think About Ethen Robotics

    The Ethen Robotics Research Lab is a research direction, not a hardware program. It studies the problems every acting agent faces — pursuing goals over many steps, understanding the state of its environment, predicting what its actions will do, recovering when they go wrong, knowing when to stop, and proving that a task is done — and it studies them in software first. Its research questions run from near-term work on reliable action in software, through skills that survive changes in models and tools and digital state models that predict effects before acting, to the far-off question of grounding any of this in the physical world. Physical work, if it ever happens, would come through integration with existing systems, ground truth gathered with partners, and control only by qualified robotics and safety teams under recognized standards. No products, partnerships, dates or results are announced here. This article explains what the Lab studies and how we think about the path from screen to physical world.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.