Research Protocol · 2026-10-03 · Model Intelligence / Faros
A Research Protocol for Model Change Assurance
An assurance report is a prediction about live work. This protocol specifies how to find out whether those predictions are right, before anyone relies on them.
Abstract
Model change assurance proposes that a model or provider upgrade should not reach consequential work until the organization has replayed its own historical tasks under the old and new configurations, verified both, and compared them task by task. The concept, with its paired analysis and regression triage, is set out in Model Change Assurance: Testing AI Upgrades Before They Reach Real Work. An assurance report is only useful if it is predictive: the regressions it finds should appear in live work, and the families it clears should not regress live. That has not been tested. This protocol specifies the study that would test it. Historical tasks are stratified by workflow family and risk, qualified for replay fidelity, and replayed under both configurations with repeated trials. The replay produces a pre-registered, per-family regression certificate. The change is then rolled out under a randomized, bounded live comparison, and the certificate's predictions are scored against live outcomes. Primary measures are the calibration of predicted deltas and the rate of false reassurance: families certified as safe that regress live. The protocol also covers cost, latency, safety and authority regressions, verifier error, and provider-version drift. No results exist. A working title that implied first results was replaced because none have been measured.
Why the concept needs its own validation study
Replay-based assurance makes an empirical claim: that what an agent does on recorded tasks in a stubbed environment predicts what it will do on new tasks in production. Several mechanisms could break that link. Replay environments can be more forgiving than live systems, or less. Historical tasks can differ from the tasks arriving next month. Verifiers can disagree with the outcomes customers care about. Hosted models can change behind a stable name after the assessment is done (Chen et al.). The backward-compatibility literature shows that instance-level regressions, sometimes called negative flips, occur even when aggregate accuracy improves (Srivastava et al.; Yan et al.; Echterhoff et al.). It does not show that an organization can detect those regressions in advance from replay.
Without a validation study, an assurance report risks becoming a ritual: a document that looks rigorous and is filed before each switch, but whose predictions nobody has ever checked against what happened next. The aim here is to measure the predictive validity of the reports, not merely to produce them.
Research questions and hypotheses
RQ1 (predictive validity). Do per-family deltas estimated from paired replay agree with deltas observed in a randomized live comparison?
RQ2 (detection). What share of regressions that appear live were flagged by replay beforehand?
RQ3 (false reassurance). How often does a family certified as having no material regression regress live by more than the certified bound?
RQ4 (non-functional change). Do replay estimates of cost, latency, escalation and refusal changes track live changes?
Pre-registered hypotheses, each a [PROPOSED TARGET] to be fixed before data collection:
- H1. Across qualified families, the replay-estimated delta in verified success lies within the live delta's 95% interval in at least 80% of families.
- H2. Replay flags at least 70% of families that show a statistically supported live regression.
- H3. The false-reassurance rate, defined below, does not exceed the nominal error rate stated on the certificates.
These thresholds are starting points from Ethen's internal design work, not results. A study that misses them is informative: it would mean replay-based certificates should not be used as gates in their current form.
Population, families and stratification
The unit of analysis is a completed task with a recorded verified outcome, captured with enough state and tool responses to be replayed, and carrying rights that permit evaluation use. Tasks are grouped into workflow families: sets of tasks that share tools, verification method and risk profile, such as code maintenance with test suites, support case resolution with state checks, or structured data entry with schema and reconciliation checks.
Sampling is stratified by family and, within family, by risk class. Families with irreversible effects or regulatory exposure are oversampled relative to their share of traffic, because these are the families where a missed regression is most costly. Within each family, sampling also covers time, so that tasks from several months are represented, and a held-out set of the most recent tasks measures whether historical tasks still resemble current work.
Families with too few eligible tasks for a meaningful bound are reported as insufficient evidence and excluded from the confirmatory analysis. They are never reported as "no regression found".
Study design
Figure 1 shows the five stages.
Figure 1. Five stages from historical tasks to scored predictions. Families are qualified for replay fidelity, replayed in pairs with repeated trials, and given a frozen certificate before any live data exist. A bounded randomized rollout then produces the live outcomes against which each certificate is scored. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Stage A: fidelity qualification. Each family is tested for replay fidelity using the measurements described in Tenant Replay: stub coverage, and agreement between replaying the original configuration and its original verified outcome, compared with the replay-to-replay noise floor. Families below pre-registered thresholds leave the confirmatory set. The general mechanism is described in Counterfactual Replay for AI Agents.
Stage B: paired replay. Every sampled task in a qualified family is replayed under both the incumbent and the candidate configuration, with tools, context strategy, skills and verifiers held fixed. Each task is run k times under each configuration, so that reliability across trials (pass^k; Yao et al.) can be estimated and stochastic flips separated from systematic ones.
Stage C: certificate. For each family, the analysis produces a regression certificate and freezes it, with a timestamp, before any live data from the candidate are seen.
Stage D: randomized live comparison. The candidate is introduced through a bounded rollout. Eligible new tasks in each family are randomized at the task level between incumbent and candidate, following standard practice for choosing the randomization unit in online experiments (Kohavi et al.). Exposure is capped by risk class, and families with irreversible effects enter only after lower-risk families show no safety stop.
Stage E: scoring. Live outcomes are verified with the same verifier versions used in replay, and each certificate's predictions are scored against them.
What is measured
Figure 2 lists the outcome dimensions and the role each plays.
Figure 2. Outcome dimensions and their role in the decision. Verified success and authority events gate the certificate; cost, latency and behavioral changes are reported alongside it. Model-based judges never decide an authority outcome alone. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Verified task success is primary. Families use the strongest available verifier: deterministic state and invariant checks where possible, programmatic test suites where available, and calibrated model-based judges only where neither exists. Judges never decide a safety or authority outcome alone. Verifier error rates are taken from the measurement procedure in How Should We Measure the Reliability of LLM Verifiers? and propagated into intervals rather than ignored. Why verifier error limits what any outcome-based comparison can show is discussed in Evaluating the Evaluators.
Safety and authority regressions are measured separately from success. A candidate model that completes more tasks while attempting more actions outside its mandate has not improved. Authority checks are deterministic: each proposed action is compared with the scope, budget and approval rules of the governing mandate. Any attempted out-of-scope action, unapproved effect, or bypassed approval counts, whether or not the runtime blocked it. Because such events should be rare, they are reported as counts with upper confidence bounds, not rates alone.
Cost and latency are compared per verified outcome, not per call, using the unit in Cost Per Verified Outcome. A candidate that is cheaper per token but needs more attempts may cost more per verified result, a point also emphasized in analyses of agent evaluation that report cost alongside accuracy (Kapoor et al.).
Behavioral change covers escalation and approval-request rates, refusals, tool-use patterns and output length. These are reported, not gated, but they are often early signals: a rise in escalations can mean people are absorbing a capability gap before success rates show it.
Capability gains are measured with the same machinery. Positive flips are reported per family, because a decision to switch only some families requires knowing where the candidate is better as well as where it is worse.
The regression certificate
A certificate is a dated, signed statement scoped to one family, one incumbent and one candidate. It records the model and provider identifiers with any version or snapshot labels the provider exposes, the tool, skill and verifier versions, the replay fidelity measurements, the sample size and number of trials, and the estimates with intervals: delta in verified success, negative-flip count with an upper bound, change in authority events with an upper bound, and changes in cost and latency.
Its central claim takes one of three forms. No material regression means the lower confidence bound on the success delta lies above a pre-registered margin, assessed with a one-sided equivalence-style test (Schuirmann), and the authority-event bound is below its threshold. Regression means the interval excludes zero in the incumbent's favor. Inconclusive covers everything else. A certificate also states its validity conditions: it expires after a fixed period, and it is invalidated early by any of the triggers in Figure 4.
Scoring predictions against live outcomes
Figure 3 shows how certificates are scored.
Figure 3. Scoring a certificate against live outcomes. Each family's certificate is compared with its live interval. False reassurance, a family certified safe that regresses live, is the costliest error and the primary safety measure of the study. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).
Each family's certificate, compared with the live comparison, falls into one of four cells. A true alarm is a certified regression confirmed live. A false alarm is a certified regression not seen live. A true reassurance is a family certified safe that stays safe. A false reassurance is a family certified safe whose live delta falls below the certified margin. False reassurance is the costliest error, because it is the one that lets a regression reach production with a document attesting that it would not.
Scoring uses the live interval, not the point estimate, so that families with small live samples are not counted as disagreements merely because of noise. Predictive validity is also summarized continuously: the correlation and calibration slope between replay-estimated and live-estimated deltas across families, weighted by precision.
Statistical analysis
Paired replay. Within each family, discordant pairs drive the comparison of success rates (McNemar), with intervals computed from paired data. Clustering by customer, template or source system is handled with a cluster bootstrap.
Verifier error. Where a judge is used, a human-audited subsample supports a correction of the estimated success rate and its interval. Prediction-powered inference is one principled way to combine a large judged sample with a small audited one (Angelopoulos et al.).
Rare events. Zero observed authority events in n tasks bound the true rate below roughly 3/n at 95% confidence (Hanley & Lippman-Hand). Certificates report the bound, not "zero".
Multiplicity. RQ1 to RQ3 are confirmatory. Family-level tests across many families are controlled for false discovery (Benjamini & Hochberg). RQ4 and all subgroup analyses are labeled exploratory.
Power. A pilot on one family estimates discordance rates and intra-cluster correlation. Families are sized so that the certificate margin is demonstrable; families that cannot reach that size are reported as insufficient evidence.
Provider-version drift
A certificate is a claim about a specific pair of configurations. Hosted models can change behind a stable identifier (Chen et al.), and small changes in prompt formatting can move results substantially (Sclar et al.). The protocol therefore treats drift as a first-class threat. Every replay and live call records the model identifier and any version metadata the provider returns. A small, fixed anchor set of tasks from each certified family is replayed on a schedule throughout Stage D. If the anchor results move outside their noise band, the certificate is suspended and the drift is reported as a finding. The same rule applies if a tool interface or skill version changes during the study, a question developed further in How to Measure Whether AI Skills Survive a Frontier-Model Upgrade.
Figure 4. Certificate lifecycle and invalidation triggers. A certificate is valid only for the configuration pair, versions and period it names. Anchor-set drift, a tool, skill or verifier version change, a measured task-mix shift or simple expiry suspends it until it is re-certified by a new paired replay. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen research protocol (proposed).
Sources of false reassurance
Some of the ways certificates could be wrong are worth stating in advance, because the study is designed to detect them:
- Lenient stubs that credit the candidate with success on calls that would fail live.
- Task-mix shift, where new work differs from the historical sample. The recent held-out set measures this directly.
- Shared verifier blind spots, where both configurations pass tasks that are actually wrong, so the delta looks flat.
- Adaptation effects, where prompts and skills tuned to the incumbent depress the candidate's replay score, producing false alarms rather than false reassurance. Separating capability from compatibility is the purpose of the Capability Transfer Ledger.
- Selective qualification. Families that pass fidelity checks may be systematically easier, so a good validation result applies only to qualified families.
Relation to benchmark-based transfer measurement
The public counterpart of this study is VerifiedWork Transfer, which measures capability survival across model and tool changes on shared tasks. That track asks whether capability transfers in general. This protocol asks whether one organization can predict a specific upgrade's effect on its own work. The two answer different questions and are designed to be read together.
Limitations
This protocol has not been run, and its thresholds are proposed. The randomized live stage needs enough traffic per family to produce informative intervals, which many organizations will not have for their rarest and riskiest families. The study validates certificates only for families that qualify for replay, and its conclusions should not be extended to the rest. Live comparison itself exposes some work to the candidate before it is fully trusted, so it must be bounded by risk class and governed by safety stops. Results will depend on the model pair, the task families and the verifiers used, and a favorable result for one upgrade does not establish validity for the next.
Conclusion
Model change assurance is a reasonable idea whose central assumption, that replay predicts live behavior, has not been tested. This protocol states how to test it: qualify families for fidelity, freeze certificates before live data exist, compare them with a bounded randomized rollout, and count every false reassurance. If certificates are predictive, organizations gain a defensible gate for upgrades. If they are not, the study will show where and why, which is equally worth knowing.
FAQ
How is this different from Paper 11? Paper 11 proposes model change assurance and its analysis. This protocol tests whether assurance reports actually predict live outcomes.
Why not just roll out the new model and watch? Rollout without prediction gives no early warning and no way to tell which families were at risk. The study uses both, and measures how well the first anticipates the second.
What is a false reassurance? A family certified as having no material regression that regresses beyond the certified margin in live work.
Related research
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work — the concept under test.
- Tenant Replay: Private Evaluation Inside Enterprise Boundaries — tenant replay environments.
- Counterfactual Replay for AI Agents — replay engine.
- Evaluating the Evaluators: Reward Integrity for AI Agents — verifier error.
- How to Measure Whether AI Skills Survive a Frontier-Model Upgrade — skill survival under upgrades.
References
- Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
- Srivastava, M. et al. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
- Yan, S. et al. (2020). Positive-Congruent Training: Towards Regression-Free Model Updates. arXiv:2011.09161. https://arxiv.org/abs/2011.09161
- Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
- Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press. https://doi.org/10.1017/9781108653985
- Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
- Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. J. Pharmacokinet. Biopharm. 15:657–680. https://doi.org/10.1007/BF01068419
- McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2):153–157. https://doi.org/10.1007/BF02295996
- Angelopoulos, A. N. et al. (2023). Prediction-Powered Inference. arXiv:2301.09633. https://arxiv.org/abs/2301.09633
- Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
- Benjamini, Y., Hochberg, Y. (1995). Controlling the False Discovery Rate. JRSS B 57(1):289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
- Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Faros: Researching How Intelligence Should Choose Intelligence
A position paper reframing AI model routing as an execution-configuration decision across model, context, tools, verification, recovery, cost and risk.
- Why Learned AI Model Routing Must Beat Good Rules
A survey of learned LLM routing: what RouteLLM, RouterBench and LLMRouterBench show, why strong rules are the right baseline, and how to test non-inferiority.
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work
A research note on model change assurance: replaying an organization's own historical tasks to find regressions before an LLM upgrade reaches real work.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.