Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Model Intelligence / Faros

Model Change Assurance: Testing AI Upgrades Before They Reach Real Work

Publication type
Research Note
Research program
Model Intelligence / Faros
Published
Authors
Ethen Research Lab
Reading time
12 min read

A new model can be better on average and worse on exactly the work you depend on. Benchmarks will not tell you which. Your own history can.

Cover image for "Model Change Assurance: Testing AI Upgrades Before They Reach Real Work". Decorative abstract motif; contains no data.

Abstract

Organizations that build on foundation models face frequent model changes: new versions, deprecated versions, price changes that make a different model attractive, and silent behavior changes behind a stable name. Each change is a migration risk. Public benchmarks measure aggregate capability on public tasks; they do not show whether a particular organization's workflows will regress. This note proposes model change assurance: before a model change reaches live work, replay a stratified sample of the organization's own historical tasks under the current and candidate models. Verify both with the same verifiers, compare them task by task, triage each regression, and produce a change report with per-family deltas and confidence intervals. We ground the proposal in the backward-compatibility literature, which shows that model updates can introduce instance-level regressions, called negative flips, even when aggregate performance improves. We discuss paired statistical analysis, the triage needed to separate real regressions from noise, verifier artifacts and replay infidelity, and the conditions under which an assurance report can be trusted. LLM upgrade regression testing of this kind is an Ethen architecture proposal; no assurance result is reported here.

Why model changes are a migration problem

A model is usually chosen once and then changed repeatedly. Providers release new versions and retire old ones. Prices fall, and a cheaper model becomes attractive. Behavior can also change behind a stable model name: a study of a widely used hosted model found substantial differences in behavior between versions released months apart on several tasks (Chen et al.). Prompt behavior is sensitive to formatting choices that carry no meaning, so a prompt tuned for one model can perform differently on another (Sclar et al.).

For an organization running agents on real work, each change raises the same questions. Which of our workflows will get better, which will get worse, and how will we know before customers do? Aggregate benchmark scores do not answer this, even when they are reported across many scenarios and metrics, as holistic evaluations do (Liang et al.). A new model that improves average accuracy by several points can still fail on a specific class of tasks that the previous model handled reliably. Those failures may concentrate in exactly the workflows an organization has tuned most carefully.

What the backward-compatibility literature shows

Machine learning research has studied this problem under the name backward compatibility. In image classification, Yan et al. define negative flips: test samples that a new model misclassifies but the old model classified correctly. They show that reducing negative flips is a distinct objective from reducing overall error, and propose training methods that target it. Srivastava et al. study how updates intended to improve models can introduce new errors that affect downstream systems and users, across architectures and settings including data shifts and inferential pipelines. For large language models, Echterhoff et al. observe model update regression: when a pretrained base model is updated, previously correct instances of fine-tuned downstream adapters become incorrect. They propose an update strategy to reduce such flips.

Two lessons carry over to agents. First, instance-level change is the right unit of analysis, not aggregate accuracy. Second, regressions are expected, not exceptional. An assurance process must find them rather than assume they will be rare.

Proposal: assurance on your own work

Model change assurance applies these lessons to deployed agent work (Figure 1).

Pipeline left to right: trigger (new model or provider version, price change, deprecation notice); task selection (stratified historical tasks including rare and high-risk families); safe replay of each task under the current and candidate model with identical tools and verifiers; paired verification; regression triage; change report with per-family deltas and confidence intervals; decision (switch, partial switch by family, hold); post-switch monitoring feeding back into task selection.

Figure 1. The model change assurance pipeline. A candidate model is evaluated on the organization's own historical tasks, replayed safely, before it reaches live work. The output is a change report, not a single score: which task families improve, which regress, why, and with what confidence. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (Model Change Assurance).

Trigger. A new model or provider version, a price change, a deprecation notice, or a scheduled re-evaluation.

Task selection. A stratified sample of the organization's historical tasks, drawn from receipts that record each task's verified outcome. Stratification ensures that rare and high-risk task families appear in proportion to their importance, not their frequency. A sample drawn uniformly from recent traffic will be dominated by easy, common work.

Safe replay. Each selected task is replayed under the current model and the candidate model, with identical tools, context strategy and verifiers, and with side effects stubbed. The general machinery and its fidelity limits are described in Counterfactual Replay for AI Agents. The in-boundary deployment for customer data is described in Tenant Replay.

Paired verification. Both runs are judged by the same verifiers, chosen for each family according to the reliability measurements in Evaluating the Evaluators.

Regression triage. Each negative flip is examined to determine whether it is a real regression or something else (see below).

Change report. Per-family deltas with confidence intervals; counts of negative and positive flips; changes in cost, latency and escalation rates; a list of families where replay could not be trusted; and examples of each confirmed regression.

Decision. Switch fully, switch for some task families only, or hold. Partial switching is often the right answer: a model can be better for research synthesis and worse for structured data entry.

Post-switch monitoring. Live outcomes after the switch are compared with the assurance report's predictions. Discrepancies feed back into task selection and into validation of the replay itself.

Paired outcomes, not averages

The core analytical object is the paired outcome table (Figure 2).

Two-by-two grid of paired outcomes. Rows: current model passes, current model fails. Columns: candidate passes, candidate fails. Cells: both pass (stable); current pass and candidate fail (negative flip, highlighted, regression); current fail and candidate pass (positive flip, improvement); both fail (persistent failure). A note says aggregate accuracy can rise while negative flips remain numerous.

Figure 2. Paired outcomes reveal what averages hide. Each historical task falls into one of four cells. Aggregate accuracy compares only the margins; assurance depends on the off-diagonal cells, especially negative flips: tasks the current model handled that the candidate fails. Evidence label: CONCEPTUAL DIAGRAM. Source: Concept from backward-compatibility literature (negative flips); illustrative.

Every replayed task falls into one of four cells: stable success, negative flip, positive flip, or persistent failure. Aggregate accuracy compares only the row and column totals. Assurance depends on the off-diagonal cells. Paired binary outcomes call for paired statistics. McNemar's test uses exactly the discordant pairs to test whether the two models differ (McNemar), and confidence intervals for the difference in success rates should likewise be computed from paired data. Pairing makes assurance far more efficient than comparing two independent samples, because the shared difficulty of each task cancels out. The same logic underlies variance-reduction practice in online experimentation (Kohavi et al.).

The report should present negative flips as a count with examples, not only as a rate. An organization may accept a net improvement of five points but not a single negative flip on a regulated workflow. That judgment belongs to the organization, and the report should give it the information to make it.

Triage: not every flip is a regression

A negative flip has at least five possible explanations (Figure 3).

Table of five explanations for a negative flip: true capability regression, stochastic variation, verifier artifact, prompt or format incompatibility, replay infidelity. Columns: distinguishing evidence and next action. For example stochastic variation is identified by repeated runs of both models; verifier artifact by expert review of the flipped case.

Figure 3. Triage: is a negative flip a real regression?. Before a negative flip counts as a regression, competing explanations must be ruled out. The evidence column lists what distinguishes each explanation; the action column lists what follows. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen methods proposal.

Stochastic variation. Agents are not deterministic. A task that passes once and fails once under the same model is not evidence of a regression. Running each task several times under both models, and reporting reliability measures such as pass^k (the probability of succeeding on all k independent trials; Yao et al.), separates variation from change.

Verifier artifacts. The candidate may produce a correct output in a form the verifier does not recognize. Expert review of a sample of flips estimates how often this happens.

Prompt or format incompatibility. The candidate may fail because prompts, tool descriptions or skills were tuned for the current model. Such failures are a migration task, not a capability loss, and often disappear after adapting the prompt or skill. Separating capability from compatibility is the purpose of the Capability Transfer Ledger.

Replay infidelity. If the candidate takes actions the original run never took, and the replay environment responds unrealistically, the flip may be an artifact of the environment.

True regression. What remains after the other explanations are excluded.

How many tasks are enough?

The sample size an assurance report needs depends on the smallest regression rate the organization cares about. A useful rule of thumb comes from the statistics of rare events. If a task family shows zero negative flips in n independent tasks, the upper 95% confidence bound on its true flip rate is approximately 3/n (Hanley & Lippman-Hand). [ILLUSTRATIVE EXAMPLE — arithmetic, not an Ethen result.] Zero flips in 100 tasks bounds the flip rate below about 3%; bounding it below 1% needs roughly 300 flip-free tasks in that family. Clustering of tasks, for example many tasks from one customer or template, widens these bounds further.

Two practical consequences follow. First, an organization should decide in advance which families need tight bounds, usually those with irreversible effects or regulatory exposure, and sample them heavily, even if they are rare in traffic. Second, families with too few historical tasks to support a meaningful bound should be reported as insufficient evidence, not as no regressions found. A report that says "no regressions observed" without the sample size behind it invites exactly the false confidence assurance is meant to prevent.

Beyond pass and fail

A model change can alter behavior in ways that are not failures but still matter. A candidate model may request more human approvals, which shifts work onto people. It may use more tool calls or longer reasoning, so that per-call savings disappear at the task level. It may refuse more often on borderline requests, or less often. It may produce longer outputs that downstream systems truncate. An assurance report should therefore compare, per family, not only verified success but also cost per verified outcome, latency, escalation and approval rates, refusal rates, and changes in tool-use patterns. Several of these are leading indicators: a jump in escalations may signal a capability gap that verified-success rates have not yet revealed, because humans are absorbing it. The economic comparison uses the unit defined in Cost Per Verified Outcome.

Why this is more valuable as models improve

One might expect assurance to matter less as models get better. We expect the opposite. Better and cheaper models increase the frequency of model changes, because there is more reason to switch. They also increase the stakes, because agents built on better models are given longer and more consequential work. An organization that switches models quarterly without assurance accepts an unmeasured regression risk each time. The broader argument is developed in Why Better Foundation Models May Make Evaluation More Valuable, Not Less.

Independence

Who should produce an assurance report? A model provider has detailed knowledge of its own model but an obvious interest in favorable results, and limited reason to evaluate a competitor's model on a customer's work. The organization itself has the right interest but may lack the replay infrastructure. A neutral party that runs replay inside the customer's boundary, with the customer able to pin the verifier versions used, is one answer. We do not claim neutrality is sufficient. Large cloud platforms that host many providers' models could offer similar evaluation, and the value of an independent report depends on its methods being inspectable. The safeguards in Evaluating the Evaluators apply here too.

What an assurance report cannot show

An assurance report predicts behavior on tasks like the ones replayed. It cannot show behavior on task types the organization has never run, and it cannot capture effects that replay stubs out. Its predictions are only as good as replay fidelity, verifier accuracy and the representativeness of the sample. The validity of the whole approach, whether assurance reports actually predict post-switch outcomes, is an empirical question. The protocol for answering it is A Research Protocol for Model Change Assurance.

Limitations

This note proposes a process; no assurance report has been produced or validated by Ethen. Replay fidelity for multi-system workflows is unknown. The sample sizes needed to detect regressions in rare but important task families may be large, and historical tasks may not cover them at all. Repeated runs to control for stochasticity multiply cost. Organizations may not hold enough verified historical tasks for assurance to be meaningful when they first deploy.

Conclusion

Every model change is a bet that the new model is better on the work that matters. Model change assurance replaces the bet with a measurement: replay your own work, compare task by task, triage every regression, and decide family by family. The idea is borrowed from the backward-compatibility literature; what is new is applying it to agents doing real, verified work, inside the organization's own boundary.

FAQ

Why not just rely on public benchmarks? Public benchmarks measure aggregate capability on public tasks. They do not show instance-level regressions on an organization's own workflows, which is what determines whether a switch is safe.

What is a negative flip? A task the current model handled correctly that the candidate model fails. Aggregate accuracy can rise while negative flips remain common.

Does assurance require sending data to Ethen? No. The proposed design runs replay inside the customer's boundary, with only aggregate results leaving where permitted.

References

  1. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  2. Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
  3. Yan, S. et al. (2020). Positive-Congruent Training: Towards Regression-Free Model Updates. arXiv:2011.09161. https://arxiv.org/abs/2011.09161
  4. Srivastava, M., Nushi, B., Kamar, E., Shah, S., Horvitz, E. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. KDD 2020. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
  5. Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
  6. Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
  7. McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2):153–157. https://doi.org/10.1007/BF02295996
  8. Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
  9. Liang, P. et al. (2022). Holistic Evaluation of Language Models. arXiv:2211.09110. https://arxiv.org/abs/2211.09110
  10. Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press. https://doi.org/10.1017/9781108653985

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Product

    How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading Paths

    The fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.

  • Infrastructure

    Why Ethen Supports More Than One AI Provider

    Ethen supports more than one AI provider because no single provider is the best fit for every job, available at every moment, acceptable under every data agreement, or stable in price and behavior over time. Using several providers gives Ethen resilience when one is degraded, a choice of the right model for each kind of work, options that meet different data-residency and contractual rules, and room to respond when models and prices change. Multi-provider support also has real costs: models behave differently across providers, there are more combinations to test, and fallback paths fail when they are rarely exercised. Ethen manages those costs with a few rules — check eligibility before preference, never break a data or contract rule to stay available, keep the state of a task independent of any one provider, and record which provider served each request.

  • Models & Intelligence

    What We Look for Before Adding a New Model to Ethen

    Before adopting a new AI model, Ethen evaluates it against ten questions. Do we know exactly which model and version it is? Does it fill a gap — a medium, a task, a price or speed point, a local or open option — that Ethen cannot already fill well? Does it perform well on tasks like the ones people actually bring to Ethen, not only on public benchmarks? Is it reliable under real conditions? Will it change without notice? Do its data-handling terms and license allow the uses our users need? How does it behave with risky requests and untrusted content? Is its pricing clear and dated? Does it fit how Ethen calls models and tools? And where, if anywhere, should people see it? A model can be listed in our catalog with sourced facts long before it is qualified for a particular use, and qualified long before it is surfaced in a curated place like Ethen Chat. Each step requires its own evidence.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.