Skip to content

EthenEthenEthen

Why Ethen Sometimes Won’t Use the Newest Model

Should you upgrade to the newest AI model? Not automatically. Ethen sometimes holds back from the newest model because a model that is better on average can still be worse on the specific work people rely on — and averages hide exactly those regressions. A new version may follow formats differently, refuse different requests, call tools differently, cost more per finished task, run slower, or come with different terms. Prompts and skills tuned for the previous model may perform worse until they are adapted. So Ethen treats a new model as a candidate, not an upgrade: we test it on representative work task by task, look past the average to regressions, cost and behavior, switch kind of work by kind of work where it wins, keep tested versions pinned and a fallback available, and keep watching after a switch. Sometimes that means adopting a new model within days. Sometimes it means not adopting it at all.

Should you upgrade to the newest AI model? Not automatically. Ethen sometimes holds back from the newest model because a model that is better on average can still be worse on the specific work people rely on — and averages hide exactly those regressions. A new version may follow formats differently, refuse different requests, call tools differently, cost more per finished task, run slower, or come with different terms. Prompts and skills tuned for the previous model may perform worse until they are adapted. So Ethen treats a new model as a candidate, not an upgrade: we test it on representative work task by task, look past the average to regressions, cost and behavior, switch kind of work by kind of work where it wins, keep tested versions pinned and a fallback available, and keep watching after a switch. Sometimes that means adopting a new model within days. Sometimes it means not adopting it at all.

Key takeaways

  • Newest is not the same as best for your work. Average improvements can hide specific regressions.
  • Compare task by task. The tasks that flip from success to failure matter more than the overall score.
  • Behavior changes count, not just accuracy. Formats, refusals, tool use, verbosity, cost and speed all shift.
  • Switch where it wins. A model can be better for research and worse for data entry; adopt it where the evidence supports it.
  • Pin, fall back, watch. Tested versions stay pinned, a fallback stays available, and switches are monitored.

Why isn't the newest model always better?

The newest model is not always better for a given workflow because model quality is not one number. A new release is typically better on average across the benchmarks its developer cares about. That says little about whether it is better on your work, under your prompts, tools and formats.

Machine learning research has a name for the problem: backward compatibility. Studies of model updates found that a new model can misclassify examples the old model handled correctly — so-called negative flips — even when its overall error is lower, and that reducing those regressions is a different goal from improving the average (Yan et al., 2020). A broader analysis showed how updates intended to improve models can introduce new errors that affect downstream systems and users (Srivastava et al., 2020). For large language models specifically, researchers have observed that updating a base model can turn previously correct outputs of systems built on top of it into incorrect ones (Echterhoff et al., 2024).

Two-by-two table: old model succeeded or failed versus new model succeeds or fails; cells are stable success, negative flip, positive flip, persistent failure.
Figure 1. A new model can raise the average while breaking specific tasks. Only a task-by-task comparison shows which.

What else changes when a model changes?

Beyond accuracy, a new model version can change at least six things that workflows depend on.

Formats. Output length, structure and adherence to required formats can shift. A downstream system expecting a particular structure can break even when the content is better.

Prompt sensitivity. Prompts tuned for one model may perform differently on another. Research has shown that model performance can change substantially with formatting choices that carry no meaning to a human reader (Sclar et al., 2023). A prompt library built over a year can degrade silently with one upgrade.

Refusals and caution. A new model may decline requests the old one handled, or handle requests the old one declined. Either can matter.

Tool use. Changes in how a model calls tools, handles errors from tools, or decides when to stop can break agentic workflows that worked reliably.

Cost and speed. A model that reasons longer may produce better answers at higher cost and latency per task. Cheaper tokens do not guarantee cheaper finished work if more retries or corrections are needed.

Behavior behind a stable name. Hosted models can change without a visible version change. A study of a widely used hosted model documented substantial behavior differences between versions released months apart (Chen et al., 2023). That is why we prefer models whose versions can be pinned.

How does Ethen decide whether to adopt a new model?

Ethen decides whether to adopt a new model by testing it on evidence, not on announcements.

Five boxes: Pin what works, Test on real work, Look past the average, Switch where it wins, Watch after switching.
Figure 2. Upgrading is a decision made on evidence, family by family — sometimes quickly, sometimes not at all.

Pin what works. Workflows that depend on consistent behavior keep the version they were tested with until a change has been checked.

Test on real work. The candidate runs on the same representative tasks as the current model, under the same tools and checks, and results are compared one task at a time. Paired comparisons like this are statistically efficient because each task's difficulty cancels out; the classic test for paired success-or-failure outcomes uses exactly the tasks where the two models disagree (McNemar, 1947).

Look past the average. We count regressions — tasks that used to succeed and now fail — and look at cost, speed, refusals, formats and tool behavior, not just overall quality. A few regressions on a workflow with irreversible actions can outweigh a large average gain elsewhere.

Switch where it wins. Adoption happens by kind of work. A model may become the default for research summaries while the previous model stays the default for structured extraction.

Watch after switching. A fallback stays available, and live results are compared with what testing predicted. If they diverge, the decision is revisited.

Ethen Research Lab describes a more systematic version of this practice as Model Change Assurance: replay a stratified sample of an organization's own past tasks under the old and new models, verify both, and triage every regression to distinguish real capability loss from noise, format incompatibility or test artifacts. It is a research note describing an architecture proposal; a separate protocol designed to test whether such assurance actually predicts live outcomes has not been run.

What about skills and prompts built for the old model?

Skills, prompts and procedures built for one model are a hidden part of every upgrade. When a new model arrives, a skill can keep helping, help less, stop mattering because the new model no longer needs it, or start hurting by constraining a capable model to an old workaround. That last case — negative transfer — is the easiest to miss, because nobody checks how the new model performs without the old instructions.

Ethen Research Lab has proposed measuring this directly: run each skill with and without it, on the old and the new model, and classify the result as transferred, degraded, obsolete or negatively transferred. The method is described in The Capability Transfer Ledger, and a protocol for doing it around a frontier-model upgrade is described in How to Measure Whether AI Skills Survive a Frontier-Model Upgrade. Both are research — a methods paper and a protocol — with no measurements yet. The practical lesson is available now: when a model changes, re-test the instructions you built for the old one, and be willing to retire the ones the new model no longer needs.

When does Ethen adopt a new model quickly?

Ethen adopts a new model quickly when it fills a clear gap and tests well. If a new model handles a medium or task that existing options handle poorly, or offers a large improvement on representative tasks without regressions where it matters, waiting has a cost too. The point of the process is not caution for its own sake; it is to make the decision on evidence. Sometimes the evidence arrives in days.

The criteria that come before any of this — whether a model belongs in Ethen at all — are in What We Look for Before Adding a New Model to Ethen.

What should you do with your own upgrades?

You can apply the same discipline without special tools.

  • Keep a test set. Twenty to fifty real examples of each important workflow, with pass criteria written down.
  • Pin versions where you can. Know which version each workflow uses.
  • Compare task by task. Run old and new on the same examples; list every task that got worse, not just the overall score.
  • Check the non-accuracy changes. Cost per good result, speed, formats, refusals, tool behavior.
  • Switch gradually. One workflow at a time, with the old model available as a fallback.
  • Re-test your prompts. Remove instructions the new model no longer needs; fix those it misreads.

Our broader guide to model choice is Making AI Model Choice Less Confusing.

The cost side of an upgrade

Upgrades are often described as free improvements. They rarely are. A newer model may produce better answers by reasoning for longer, which raises both the cost and the time per task. It may be more verbose, which costs more and can break length limits downstream. It may be priced differently per token in ways that change which workflows are affordable. Or it may be cheaper per token but need more retries on some kinds of work, so that the cost of each good result goes up.

That is why we compare models on the cost of results that pass their checks, not on price per token. Ethen Research Lab's note Cost Per Verified Outcome proposes this measure for agent work: all the costs of producing results, including failed attempts and checking, divided by the number of results independently verified as successful. It is a research synthesis; Ethen's own figure has not been measured. A model that is better and more expensive per good result may still be the right choice for high-stakes work and the wrong choice for routine work — another reason to switch by kind of work rather than everywhere at once.

Questions to ask any AI product about model upgrades

If you rely on an AI product — Ethen or anyone else — these questions reveal how it handles model change.

  1. Which model version is my workflow using, and can I see it?
  2. Will the product switch my workflow to a new model automatically? If so, is it tested first, and can I opt out?
  3. Can I pin a version for work that needs consistent behavior?
  4. How are regressions detected? Ask whether new versions are compared task by task on representative work, or only on public benchmarks.
  5. What happens to my prompts, skills and settings when the model changes?
  6. Is there a fallback if the new version misbehaves?

A product that can answer these clearly treats your workflows as something to protect. A product that cannot is asking you to discover regressions yourself.

Why does this matter more as models improve?

This matters more as models improve because better and cheaper models give everyone more reasons to switch, more often — and each switch is a chance for something to break. Ethen Research Lab's position paper Why Better Foundation Models May Make Evaluation More Valuable, Not Less argues that frequent model change is one of the main reasons the ability to test changes on your own work becomes more valuable over time. It is a research synthesis and argument, with stated conditions under which it would be wrong — for example, if model updates became reliably backward compatible on real workloads.

An example

Illustrative example — hypothetical.

A new version of a model a team uses for customer-support drafts is released, with strong benchmark gains. The team runs 200 past tickets through both versions. Overall, the new version's drafts are rated better more often. But 11 tickets that the old version handled correctly now fail: in each, the new version promises a refund timeline the company does not offer, apparently because it is more eager to reassure. The new version is also noticeably more verbose, pushing some replies past the help desk's character limit. The team adopts the new version for internal ticket summaries, where it is clearly better and the risks do not apply, keeps the old version for customer-facing drafts, adjusts the prompt to state refund policy explicitly, and re-tests. Two weeks later, the re-test shows the regressions gone, and the switch is extended to customer drafts.

Tradeoffs

Testing before switching means sometimes using an older model for a while after a better one exists, and some improvements arrive later than they could. Keeping several versions available adds maintenance. Task-by-task comparison takes effort and needs a representative set of tasks. We think these costs are smaller than the cost of discovering regressions in front of customers.

Frequently asked questions

Should I always upgrade to the latest LLM? Not automatically. Test it on your own work first, compare task by task, and switch where it is better.

Why did my prompts break after a model update? Models respond differently to the same instructions, and small formatting choices can change results. Re-test prompts after an update, and remove instructions the new model no longer needs.

Does Ethen ever adopt new models quickly? Yes, when a model fills a clear gap and tests well on representative work.

What is a negative flip? A task the old model handled correctly that the new model fails. Averages can improve while negative flips remain.

References

  1. Yan, S. et al. (2020). Positive-Congruent Training: Towards Regression-Free Model Updates. arXiv:2011.09161. https://arxiv.org/abs/2011.09161
  2. Srivastava, M. et al. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
  3. Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
  4. Sclar, M. et al. (2023). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. arXiv:2310.11324. https://arxiv.org/abs/2310.11324
  5. Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
  6. McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2):153–157. https://doi.org/10.1007/BF02295996
  7. Ethen Research Lab (2026). Model Change Assurance. Research note; architecture proposal. https://upcube.ai/resources/research/model-change-assurance
  8. Ethen Research Lab (2026). The Capability Transfer Ledger. Methods paper; no measurements. https://upcube.ai/resources/research/capability-transfer-ledger
  9. Ethen Research Lab (2026). How to Measure Whether AI Skills Survive a Frontier-Model Upgrade. Research protocol; not yet run. https://upcube.ai/resources/research/skill-survival-protocol