Position Paper · 2026-10-03 · Model Intelligence / Faros
Why Better Foundation Models May Make Evaluation More Valuable, Not Less
If models become ten times better and ten times cheaper, much of what AI companies build around them loses value. The ability to know what an agent actually did, and whether it worked, is likely to gain.
Abstract
Any organization building on foundation models must ask what remains valuable if those models become dramatically better and cheaper. This position paper applies a stress test with two scenarios, models ten times more capable and inference ten times cheaper, to the components of an agent business. We argue that prompt and harness tricks, price-arbitrage routing and many narrow owned models lose value under one or both scenarios. By contrast, AI evaluation in a specific sense becomes more valuable: calibrated verifiers, private and fresh evaluation sets, verified outcome records, and assurance that a model change will not break existing work. We identify four mechanisms. More capable models take on longer and more consequential tasks. Better and cheaper models are switched more often. Stronger optimization finds verifier blind spots. Cheaper generation makes verification a larger share of the cost of trustworthy work. For each mechanism we state what would falsify it, and we take the strongest counterarguments seriously: self-verifying models, providers bundling evaluation, and commoditized evaluation tooling. The paper is an Ethen Research Lab position. It reports no measured results.
The stress test
Strategy around foundation models has a recurring failure mode: building value on a gap that the next model release closes. Prompt engineering that compensates for a model's weaknesses, routing that saves money by exploiting large price differences, and narrow models trained to beat a general model on one task are all bets that today's gaps persist. Some persist. Many do not.
A useful discipline is to imagine two extreme but plausible futures and ask what each component of a system is worth in them:
- Ten times more capable models. Agents complete much longer tasks, make fewer errors per step, and handle work that today requires specialized tooling.
- Ten times cheaper inference. The marginal cost of generation falls by an order of magnitude, and with it the savings from choosing a cheaper model.
Figure 1 summarizes our assessment. It is an argument, not a measurement, and the rest of this paper explains it.
Figure 1. Stress test: what gains and loses value. A qualitative stress test of capabilities an AI organization might build, under two scenarios: models ten times more capable, and inference ten times cheaper. The judgments are arguments made in this paper, not measurements. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis (stress test); argument, not measurement.
What loses value
Prompt and harness tricks. Much agent engineering compensates for model weaknesses: elaborate instructions, rigid output formats, step-by-step scaffolding, retry heuristics. A more capable model often does not need them. Worse, some become harmful, constraining a capable model to an older model's workaround. Practitioner experience inside the Ethen research corpus describes harness techniques as having a half-life of months, which is consistent with this. Such engineering is necessary to compete today and is not a durable asset.
Routing margin. Routing saves money when the gap between expensive and cheap models is large and many requests are easy. If inference becomes ten times cheaper across the board, the absolute savings shrink toward zero. Published evidence already shows that learned routers often fail to beat simple baselines (Li et al.), which makes routing margin an even weaker foundation. What may survive is knowledge of which configuration, meaning context, tools, verification and recovery, produces verified work, as argued in Faros. That is not a price arbitrage. The case that learned routing must beat good rules to matter at all is made in Why Learned AI Model Routing Must Beat Good Rules.
Narrow owned models. A small model trained to match a frontier model on one task is valuable while the frontier model is expensive and the small one is cheaper. Cheaper inference moves the crossover point, and a more capable frontier model may surpass the specialist. Specialists remain valuable where they encode data others lack, but they need refreshing, and their value is partly a timing bet.
What gains value, and why
Mechanism 1: longer tasks raise the cost of undetected errors
Capability gains translate into longer and more autonomous tasks. One measure is the length of task, in human time, that an AI system can complete with 50% success. By that measure the frontier has been doubling roughly every seven months since 2019, according to one analysis, which attributes the growth mainly to greater reliability, adaptation to mistakes, reasoning and tool use [EXTERNAL PRIMARY-SOURCE RESULT] (Kwa et al.). As tasks lengthen, an undetected error does more damage: it propagates through more steps, touches more systems and is discovered later. A model that is ten times better but is trusted with tasks a hundred times longer creates more need for verification, not less.
Falsifier. If error rates fall faster than task length and consequence grow, so that cheap spot checks catch nearly all problems, this mechanism weakens.
Mechanism 2: more frequent switching creates migration risk
Better and cheaper models give organizations more reasons to switch: a new version, a cheaper provider, a model better at one family of tasks. Each switch is a migration with regression risk. Model updates in machine learning systems can introduce instance-level regressions even when aggregate accuracy improves (Yan et al.; Srivastava et al.; Echterhoff et al.). Hosted model behavior can also shift between versions (Chen et al.). The more often an organization switches, the more often it needs to know which of its workflows will break. That is the purpose of Model Change Assurance.
Falsifier. If model updates become reliably backward compatible on real workloads, so that regressions vanish in practice, assurance demand falls.
Mechanism 3: stronger optimization finds verifier blind spots
When agents are trained or tuned against verifiers, more capable optimizers find more ways to satisfy the verifier without satisfying the intent. Studies of reward misspecification found that more capable agents often exploit misspecified rewards more, sometimes with abrupt shifts in behavior as capability crosses thresholds (Pan et al.). Optimizing against a learned reward model first improves and then degrades performance on the true objective as optimization pressure increases (Gao et al.). Better models therefore make it more important to know how often each verifier is wrong and where. The program for measuring this is Evaluating the Evaluators.
Falsifier. If future training methods become robust to verifier error, or verifiers become so accurate that exploitation is negligible, this mechanism weakens.
Mechanism 4: cheaper generation makes verification a larger share of cost
Producing an answer and establishing that it is correct are different costs. Generation gets cheaper with every model price cut. Establishing correctness may not: running test suites, checking state across systems, expert review of judgment-heavy outputs and human approval of consequential actions do not become ten times cheaper when tokens do. If generation becomes much cheaper while verification does not, verification becomes a larger share of what a trustworthy outcome costs (Figure 3, below). Cheap, reliable verification then becomes a larger lever on total cost, and a more defensible capability. The economic unit that makes this visible is Cost Per Verified Outcome.
Falsifier. If verification itself becomes as cheap as generation, for example because models verify each other reliably at negligible cost, this mechanism weakens.
Figure 2 summarizes the four mechanisms.
Figure 2. Four mechanisms linking better models to evaluation demand. Each mechanism is an argument with an identifiable falsifier, discussed in the text. Together they suggest demand for trustworthy evaluation grows with capability rather than shrinking. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis; argument.
Figure 3. Cost composition of a trustworthy outcome, schematically. Schematic only, not to scale and not measured: if generation gets much cheaper while the work of establishing that an outcome is correct does not, verification becomes a larger share of what a trustworthy outcome costs. Evidence label: ILLUSTRATIVE — NOT MEASURED ETHEN DATA. Source: Schematic illustration; no data.
Public benchmarks age quickly
A related argument concerns where evaluation value lies. Public benchmarks lose value as models are trained on them or tuned to them. A carefully matched replacement for a widely used math benchmark revealed accuracy drops of up to 8%, with several model families showing evidence of systematic overfitting [EXTERNAL PRIMARY-SOURCE RESULT] (Zhang et al.). Benchmarks built from frequently refreshed sources try to stay ahead of contamination (White et al.; Jain et al.). Better models trained on more data intensify this dynamic. Evaluation value therefore shifts toward private, fresh and task-specific evaluation: an organization's own work, held out and verified. That kind of evaluation cannot be absorbed into a model's training data, because it is not public. Which data properties survive model improvements is discussed in What Makes AI Data Defensible?.
A test we can run
The thesis makes a prediction that can be checked within one model generation. When a materially better model is released, an organization that has calibrated verifiers and verified history should be able to answer three questions quickly and accurately. Which of our workflows improve? Which regress? What does a verified outcome now cost? An organization without them should find those answers slow, partial or wrong, and should discover regressions in production. If both kinds of organization adapt equally well, because the new model is good enough that nothing important breaks and nothing needs checking, the thesis is weakened. The comparison is not hypothetical. Every major model release is a natural experiment, and the protocol in A Research Protocol for Model Change Assurance describes how to measure whether assurance predicted what actually happened.
A second prediction concerns cost composition. If verification becomes a larger share of the cost of trustworthy work as inference prices fall, organizations should report rising verification spend relative to generation spend, holding verified outcome volume constant. That ratio is measurable from task-level cost records, and it is one of the quantities we intend to track.
Counterarguments
Models will verify themselves. Better models may check their own work more reliably. Self-correction without external feedback has been weak in reasoning tasks (Huang et al.), but that could change. Even if it does, an organization that relies on a model's self-assessment still needs evidence that the self-assessment is calibrated on its own work. That evidence is itself an evaluation.
Providers will bundle evaluation. Model providers and cloud platforms can and do offer evaluation tools. This is a real competitive threat to any independent evaluation business. Two considerations temper it. A provider has limited incentive to certify a competitor's model on a customer's work, and buyers may discount self-certification. And the most valuable evaluations are built on the customer's private work, which favors whoever can operate inside the customer's boundary. Neither consideration secures an independent evaluator's position.
Evaluation tooling is commoditizing. Evaluation frameworks are increasingly open and standardized. We agree. The argument here concerns calibrated verifiers, private task sets, outcome records and assurance processes, which are expensive to build and slow to accumulate. Tooling is a small part of that.
Many tasks will become trivially reliable. For simple tasks, a sufficiently good model may be reliable enough that verification is unnecessary. That is plausible and welcome. Evaluation value then concentrates on the frontier of difficulty and consequence, which moves outward with capability rather than disappearing.
What this implies
If the argument holds, organizations building on foundation models should invest in components that gain value with model progress. Those components are calibrated verifiers, private and fresh evaluation sets, verified outcome records, assurance for model changes, and a clean history of rights. They should treat prompt tricks, routing margins and narrow distillations as cost optimizations with limited lifetimes. Benchmark design for this purpose is taken up in Ethen VerifiedWork.
Limitations
This is an argument, not a measurement. The ten-times scenarios are deliberately extreme. Each mechanism has a stated falsifier, and some may turn out to apply. The external results cited concern specific settings; for example, the time-horizon analysis reports its own limits on external validity. The commercial implications depend on competitive dynamics that this paper does not model, and no Ethen measurement supports or refutes the thesis yet.
Conclusion
Better and cheaper models erode the value of compensating for model weaknesses and of arbitraging model prices. They increase the value of knowing what an agent did, whether it worked, and whether the next model will still do it. The more capable and autonomous AI becomes, the more the scarce input is trustworthy measurement rather than intelligence.
FAQ
Doesn't a better model need less checking? Per step, perhaps. But better models are given longer, more consequential tasks and are switched more often, and optimizers find verifier blind spots more effectively. Total demand for trustworthy verification can rise.
Will public benchmarks remain useful? For comparing models broadly, yes. For deciding whether a model will work on a specific organization's tasks, private and fresh evaluation on that work becomes more important as public benchmarks are absorbed into training data.
What would show this thesis is wrong? Reliable self-verification at negligible cost, backward-compatible model updates on real workloads, or verification becoming as cheap as generation.
Related research
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work — assurance demand rises with model churn.
- Evaluating the Evaluators: Reward Integrity for AI Agents — verification share of cost.
- What Makes AI Data Defensible? — which data survives.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI — economics under cheap inference.
- Why Learned AI Model Routing Must Beat Good Rules — routing margin decays.
- Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action — what to measure.
References
- Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499
- Li, H. et al. (2026). LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. arXiv:2601.07206. https://arxiv.org/abs/2601.07206
- Yan, S. et al. (2020). Positive-Congruent Training: Towards Regression-Free Model Updates. arXiv:2011.09161. https://arxiv.org/abs/2011.09161
- Srivastava, M. et al. (2020). An Empirical Analysis of Backward Compatibility in Machine Learning Systems. arXiv:2008.04572. https://arxiv.org/abs/2008.04572
- Echterhoff, J. et al. (2024). MUSCLE: A Model Update Strategy for Compatible LLM Evolution. arXiv:2407.09435. https://arxiv.org/abs/2407.09435
- Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
- Pan, A. et al. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. arXiv:2201.03544. https://arxiv.org/abs/2201.03544
- Gao, L., Schulman, J., Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760. https://arxiv.org/abs/2210.10760
- Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332
- White, C. et al. (2024). LiveBench: A Challenging, Contamination-Limited LLM Benchmark. arXiv:2406.19314. https://arxiv.org/abs/2406.19314
- Jain, N. et al. (2024). LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. arXiv:2403.07974. https://arxiv.org/abs/2403.07974
- Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv:2310.01798. https://arxiv.org/abs/2310.01798
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Work Receipts: A Verifiable Record for Autonomous AI Work
A technical report proposing the Work Receipt: one signed record of authority, actions, effects, verification, cost and rights for every unit of autonomous AI work.
- Evaluating the Evaluators: Reward Integrity for AI Agents
A methods paper on reward integrity for AI agents: verifier false accepts and rejects, abstention, grader drift, expert disagreement and reward hacking.
- Counterfactual Replay for AI Agents
A research proposal for counterfactual evaluation of AI agents: replaying completed tasks under alternative models, tools, context and recovery strategies.
Explained on the Ethen Blog
- How to Read an AI Model Comparison
Most model comparisons answer a narrower question than their headline suggests. Here is how to find the real question, check whether the numbers are comparable, and read the gaps.
- Why Ethen Is Investing in Model Intelligence
Ethen invests in model intelligence because every decision about which AI model to use — made by a person, by Ethen's AI Gateway, or by Ethen's automatic model choice — is only as good as the facts behind it. Model intelligence is the knowledge needed to make that decision: what a model can do, what it costs, how it performs on which kinds of task, and where each of those facts came from. The model landscape changes too quickly, and public comparisons hide too much, for that knowledge to be assembled ad hoc. So Ethen builds it deliberately: every fact with a source and an owner, Unknown shown instead of guesses, eligibility decided before preference, and a long-term aim of connecting model choices to whether the resulting work actually succeeded.
- What “Done” Should Mean for an AI Agent
For an AI agent, "done" should mean that every requirement of the task has been met and that something other than the agent's own report shows it. Precisely: a task is complete when each required obligation is supported by evidence at the level of checking it needs — a passing test, a reconciled record, a confirmed delivery, an approved review — or has been explicitly waived by the person who owns the task. Three refinements make the definition usable. Keep execution success (the steps ran), task success (the outcome was achieved) and business success (it produced value) apart. Treat success as provisional until it can no longer be reversed. And report partial and unknown outcomes as what they are, instead of rounding them up to done.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.