Skip to content

EthenEthenEthen

Why Better AI Models Don’t Eliminate the Need for Better Products

Better AI models do not make AI products obsolete; they change which parts of a product matter. Each model release removes some reasons to build: prompt tricks that compensated for a weaker model, savings from routing around price differences, interfaces that do little more than pass text to a model. At the same time, better models are trusted with longer and more consequential work, switched more often, and harder to check by eye — which makes the parts of a product that surround the model more valuable. Those parts are knowing what the AI actually did and whether it worked, controlling what it may touch and spend, recovering when long work is interrupted, testing that a model change will not break real work, keeping memory and data under the user's control, and interfaces built around the work rather than the chat. That is where Ethen invests.

Better AI models do not make AI products obsolete; they change which parts of a product matter. Each model release removes some reasons to build: prompt tricks that compensated for a weaker model, savings from routing around price differences, interfaces that do little more than pass text to a model. At the same time, better models are trusted with longer and more consequential work, switched more often, and harder to check by eye — which makes the parts of a product that surround the model more valuable. Those parts are knowing what the AI actually did and whether it worked, controlling what it may touch and spend, recovering when long work is interrupted, testing that a model change will not break real work, keeping memory and data under the user's control, and interfaces built around the work rather than the chat. That is where Ethen invests.

Key takeaways

  • Model releases erase workarounds, not work. Products built to compensate for model weaknesses lose value; products built around real work gain it.
  • More capable means more consequential. Better models are given longer tasks with bigger effects, which raises the cost of an unnoticed mistake.
  • Change gets more frequent. Better and cheaper models give people more reasons to switch, and every switch can break something.
  • Checking does not get cheaper at the same rate. Generating an answer gets cheaper; establishing that it is right often does not.
  • This is an argument with falsifiers. If models start reliably verifying themselves at negligible cost, the argument weakens.

Why do people think better models will make AI products obsolete?

People think better models will make AI products obsolete because they have watched it happen to a certain kind of product. When a new model can do something natively — read a long document, follow a complex format, use a tool — products whose main value was working around the old model's limits lose their reason to exist. Every major release produces a wave of commentary about which "wrappers" just died.

That observation is correct about workarounds. It is wrong about products in general, because it assumes that a product's value is the model's capability plus a thin layer. For work that matters — work that acts, persists and has consequences — most of what a person needs is not the model's capability at all. It is everything around it.

A stress test: what if models got ten times better and cheaper?

A useful way to think about this is to stress-test a product against two extreme but plausible futures: models ten times more capable, and inference ten times cheaper. Ethen Research Lab applies exactly this test in its position paper Why Better Foundation Models May Make Evaluation More Valuable, Not Less. The paper is a research synthesis and argument; it reports no measured results, and it names the evidence that would weaken its claims.

Two columns: what loses value as models improve and what gains value.
Figure 1. The argument is a stress test, not a measurement: imagine models ten times more capable and ten times cheaper, and ask what is still worth building.

What loses value. Prompt and harness tricks that make up for a model's weaknesses — elaborate scaffolding, rigid formats, retry heuristics — are often unnecessary for a stronger model, and can even get in its way. Saving money by sending easy requests to cheaper models shrinks when everything gets cheaper. Narrow tools trained to beat a general model on one task are vulnerable when the general model catches up. These are real engineering, and necessary to compete today, but they are not durable.

What gains value. The paper identifies four mechanisms that push the other way. Each is worth stating plainly, with what would prove it wrong.

1. More capable models get longer, more consequential tasks

As models improve, people give them bigger jobs. One analysis of AI systems on software tasks found that the length of task they can complete with reasonable reliability has been growing quickly (Kwa et al., 2025). A longer task has more steps where a small early error can compound, more systems it touches, and more time before anyone notices a problem. A model that is ten times better per step but trusted with tasks a hundred times longer creates more need for checking, not less.

What would weaken this: error rates falling faster than task length and consequence grow, so that cheap spot checks catch nearly everything.

2. Better models are switched more often

When better and cheaper models arrive frequently, organizations have more reasons to change: a new version, a cheaper provider, a model that is better at one kind of task. Every change is a migration with regression risk. Research on machine learning systems has found that updates can introduce new errors on specific cases even when average accuracy improves (Srivastava et al., 2020). The more often you switch, the more often you need to know which of your workflows will get better and which will break. We explain how we approach this in Why Ethen Sometimes Won't Use the Newest Model.

What would weaken this: model updates becoming reliably backward compatible on real workloads.

3. Stronger optimizers find blind spots faster

When systems are tuned against an automated check, more capable systems get better at satisfying the check without satisfying the intent. Studies of reward misspecification found that more capable agents often exploit flawed objectives more effectively (Pan et al., 2022), and optimizing hard against a learned reward model eventually makes true performance worse (Gao et al., 2022). The better the model, the more it matters that the checks around it are themselves measured and trustworthy.

What would weaken this: training methods that become robust to flawed checks, or checks so accurate that exploiting them is negligible.

4. Generating gets cheaper faster than checking

Producing an answer and establishing that it is correct are different costs. Generation gets cheaper with every price cut. Running a test suite, reconciling records across systems, having an expert review a judgment call, or getting a person to approve a consequential action does not get ten times cheaper when tokens do. As generation costs fall, checking becomes a larger share of what a trustworthy result costs — and a larger lever on total cost. Ethen Research Lab's note Cost Per Verified Outcome proposes measuring the cost of independently verified successes, failures and checking included, for exactly this reason. It is a research note; Ethen's own figure has not been measured.

What would weaken this: verification becoming as cheap as generation, for example because models check each other reliably at negligible cost.

What does a product add around a model?

A product adds four layers around a model, and better models make each of them more important, not less.

Five stacked bands: The model, Context, Workflow, Controls, Evidence.
Figure 2. A better model improves the first layer. The other four decide whether that capability can be trusted with real work.

Context. The right information at the right time — sources, files, prior decisions — and the obligations that must not be forgotten when a long conversation is compressed. A more capable model with the wrong context still produces the wrong result, faster.

Workflow. Where work lives, how it is handed off, how it survives interruptions and resumes without repeating actions. A model has no opinion about any of this; a product must. We describe why different work needs different homes in Why Ethen Is a Family of Specialized AI Apps, Not One App.

Controls. What the AI may read, change and spend, and which actions need a person's approval. These grow in importance precisely as models become capable of more.

Evidence. How results are checked and how anyone can see what happened. A model's own report of success is not evidence; a product can require evidence before work counts as done, as Ethen's mission system does. We discuss this in What "Done" Should Mean for an AI Agent.

The same model, two products

Illustrative example — hypothetical.

Imagine a new model that is excellent at reconciling invoices against purchase orders. Two products put it to work for a finance team.

The first product is a thin interface: upload files, get an answer. On easy months it is impressive. Then three things happen. The model reconciles 400 invoices and reports "all matched", but a duplicate invoice number across two subsidiaries slips through, and nobody can see which documents supported which match. A payment run triggered from the results times out, and the team cannot tell whether payments went out. Six weeks later, the provider releases a newer model; the product switches automatically, and a formatting change in the output quietly breaks the export the accounts team relies on.

The second product uses the same model, and wraps it in the layers described above. Each match links to the documents that support it, and mismatches and unknowns are listed rather than buried. Payments are prepared by the AI but require a person's approval, bound to the exact batch. When the payment run times out, the product checks the bank's records before doing anything else and finds that half the batch went out. When the newer model arrives, the team runs last month's invoices through it first, sees the export change, and fixes it before switching.

The model did not get worse in the first story or better in the second. The difference is everything around it — and every one of those differences matters more as the model is trusted with more invoices, more money and more autonomy.

How can you tell whether an AI product will survive the next model release?

Five questions are a quick test for any AI product, including ours.

  1. If the model were much better tomorrow, would this product still be needed? If its main value is compensating for today's model, probably not.
  2. Does it hold the work, or just the conversation? Products that hold state, history and artifacts gain value as tasks get longer.
  3. Does it control what the AI may do? Permissions, budgets and approvals matter more as capability grows.
  4. Can it show what happened and whether it worked? Evidence becomes more valuable as outputs become harder to check by eye.
  5. Does it make model changes safe? Products that let you test a new model on your own work before switching turn frequent releases into an advantage instead of a risk.

What are the strongest counterarguments?

The strongest counterarguments deserve a direct answer, because some of them may turn out to be right.

"Models will verify their own work." Perhaps, eventually. Today, models struggle to correct their own reasoning without external feedback (Huang et al., 2023). Even if self-checking improves, an organization relying on it still needs evidence that the self-check is accurate on its own work — which is itself a form of verification.

"Model providers will build all of this." Providers can and do build products, agents and evaluation tools, and some will be excellent. This is a genuine competitive reality. It does not change what work needs; it changes who builds it. We would rather compete on how well those needs are met than on whether they exist.

"The tooling will be commoditized." Much of it will be, and that is fine. Open standards and shared tools are good for everyone. What is harder to commoditize is getting the details right for real work: which actions need approval, how unknown outcomes are handled, how a model change is tested on an organization's own tasks.

"Some tasks will simply become reliable enough." For many simple tasks, yes — and that is welcome. The frontier of difficulty and consequence moves outward with capability; it does not disappear.

What does this mean for how Ethen builds?

It means Ethen invests most in the parts of the product that a better model makes more valuable. Concretely, that is the shared foundation across Ethen's apps — durable jobs that survive interruptions, completion that depends on evidence rather than on the agent's own report, unknown outcomes handled honestly, and approvals bound to specific actions — and a model knowledge layer that shows its sources and its gaps. We describe those foundations in What We're Building Across Ethen: October 2026 Update. It also means we treat prompt-level optimizations as necessary but temporary, and expect to revisit them with each model generation.

It does not mean we think model quality is unimportant. Better models make everything Ethen does more capable, and Ethen is built to use many of them. The point is narrower: model quality is necessary, and it is not sufficient.

Frequently asked questions

Will better LLMs make AI apps unnecessary? They make some apps unnecessary — especially those whose main value was compensating for model weaknesses. They make products built around real work, with controls, evidence and recovery, more valuable.

What value does an AI product add on top of a model? Context management, workflows that persist and recover, controls over what the AI may do, and evidence of what it did and whether it worked.

Is this argument proven? No. It is a reasoned argument with named falsifiers, set out in an Ethen Research Lab position paper that reports no measured results.

What would change Ethen's mind? Reliable self-verification at negligible cost, backward-compatible model updates on real workloads, or verification becoming as cheap as generation.

References

  1. Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499
  2. Srivastava, M. et al. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
  3. Pan, A. et al. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. arXiv:2201.03544. https://arxiv.org/abs/2201.03544
  4. Gao, L., Schulman, J., Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760. https://arxiv.org/abs/2210.10760
  5. Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
  6. Ethen Research Lab (2026). Why Better Foundation Models May Make Evaluation More Valuable, Not Less. Position paper; research synthesis. https://upcube.ai/resources/research/better-models-increase-evaluation-value
  7. Ethen Research Lab (2026). Cost Per Verified Outcome: A Better Economic Unit for Agentic AI. Research note; research synthesis. https://upcube.ai/resources/research/cost-per-verified-outcome