Skip to content

EthenEthenEthen

Research Note · 2026-10-03 · Enterprise / Sovereign AI

Cost Per Verified Outcome: A Better Economic Unit for Agentic AI

Publication type
Research Note
Research program
Enterprise / Sovereign AI
Published
Authors
Ethen Research Lab
Reading time
12 min read

Buyers of AI agents do not want tokens. They want work done, and they want to know it was done. The unit of account should say so.

Cover image for "Cost Per Verified Outcome: A Better Economic Unit for Agentic AI". Decorative abstract motif; contains no data.

Abstract

AI economics is still mostly measured in units inherited from earlier technology: cost per token, per request or per seat. For agents that perform multi-step work, these units can move in the opposite direction from the value delivered. A cheaper model can use more tokens, trigger more retries and succeed less often, so that work gets more expensive while tokens get cheaper. This note defines cost per verified outcome (CPVO): all eligible delivery costs for a defined cohort and period, divided by the number of outcomes independently verified as successful. We specify what belongs in the numerator, including failed attempts and verification itself. We specify what counts in the denominator: verified successes at the task level, with pending work reported separately. We compare CPVO with other units, relate it to recent proposals for cost-aware evaluation, and work through an illustrative calculation. We also show how verifier error biases the measured value and why outcome-based pricing built on CPVO needs safeguards against the seller grading its own work. CPVO is an Ethen research proposal. Ethen's own CPVO is unknown and has not been measured.

Why the familiar units mislead

Cost per token measures the price of a raw input. It ignores how many tokens a task needs, how often attempts fail, and whether the output was useful. Prices per token have fallen quickly, and published analyses track how the cost of reaching a given capability level has declined. For an organization deploying agents, a falling token price is necessary but not sufficient for falling costs of useful work. If inference becomes ten times cheaper, verification and human review become a larger share of the cost of a trustworthy outcome, an argument developed in Why Better Foundation Models May Make Evaluation More Valuable, Not Less.

Cost per request is closer to the work but still counts failures as if they were successes, and treats a request that resolves a problem in one call the same as one that starts a ten-step chain.

Cost per seat decouples price from cost entirely. It suits tools that augment people, and fits poorly when agents do work autonomously and usage varies by orders of magnitude between customers.

Cost per task counts attempted tasks. It includes failures in the numerator, but rewards attempting tasks rather than completing them.

Cost per successful task is closer still, but only as good as the definition of "successful". If success is self-reported by the agent, or inferred from "the run completed", the metric invites optimism, and eventually gaming.

Figure 1 compares these units on the properties that matter.

Matrix comparing six economic units (cost per token, cost per request, cost per seat, cost per task attempted, cost per successful task as self-reported, cost per verified outcome) on five properties: tracks delivered value, includes cost of failures and retries, includes verification cost, resistant to gaming, easy to measure. Cost per token is easiest to measure and tracks value least; CPVO tracks value best and is harder to measure.

Figure 1. Economic units compared. Each unit answers a different question. Only units with a verified-success denominator track what a buyer of agent work receives, and only CPVO makes the verification standard explicit. Marks are qualitative judgments. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative judgment.

Evaluation research has reached a similar conclusion from a different direction. A critique of agent benchmarks argued that a narrow focus on accuracy without attention to cost has produced needlessly complex and costly agents, and has led to mistaken conclusions about the sources of accuracy gains. It recommends optimizing accuracy and cost jointly (Kapoor et al.). An economic framework for language models defines cost-of-pass, the expected monetary cost of generating a correct solution, and tracks how the cheapest available cost-of-pass for different task types has changed over time (Erol et al.). CPVO is the counterpart of these ideas for deployed agent work, where "correct" must be established by verification against acceptance criteria, and where costs include far more than inference.

Definition

CPVO for a cohort and period = (all eligible delivery costs for that cohort and period) ÷ (number of outcomes in that cohort independently verified as successful)

Figure 2 shows the components.

Tree. CPVO for a cohort and period splits into numerator and denominator. Numerator: model and provider tokens including cache effects; tools and external APIs; sandbox and infrastructure; storage and egress; verification (tests, judges, expert review); required human review and delivery support; allocated operational overhead. Excluded from numerator and reported separately: research and training investment. Denominator: independently verified successes at task-success level; pending tasks reported separately; failures and unknowns contribute cost but not successes.

Figure 2. What goes into the numerator and the denominator. The numerator includes every eligible delivery cost for a cohort, including failed attempts and verification. The denominator counts only independently verified successes at the task-success level, with pending work reported separately rather than dropped or counted. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis (CPVO definition).

The numerator includes everything required to deliver the work. That means provider tokens with the effects of caching (cache reads, writes and misses priced as billed), tool and external API calls, sandbox and infrastructure compute, storage and egress, the cost of verification itself (running tests, model judges, expert review), required human review and delivery support, and an allocation of operational overhead. Costs of failed attempts are included, because failures are part of what it costs to produce successes. Research and model-training investment is excluded from delivery CPVO and reported separately. Hiding delivery costs in a research budget makes margins look better than they are.

The denominator counts only verified successes. A success is an outcome that met its acceptance criteria according to a named verifier, at the task-success level, after any reopen or revert window has closed. Execution success ("the calls returned") does not count. Business success ("revenue increased") is reported separately, if measurable at all, because attribution to individual tasks is weak. Three accounting rules keep the denominator honest:

  • Pending work is visible. Tasks still inside their verification window are reported as pending, not dropped and not counted. CPVO computed only on settled tasks should state how many were pending.
  • Zero is undefined, not zero. A cohort with no verified successes has undefined CPVO, not a CPVO of zero.
  • The verifier is named. Because the denominator depends on what counts as verified, every CPVO figure states the verification standard used.

CPVO should be reported with its distribution, such as median and 95th-percentile cost per task by task family, and with interval estimates. Averages across heterogeneous work hide the expensive tail where most cost risk lives.

An illustrative calculation

[ILLUSTRATIVE EXAMPLE — arithmetic only; not an Ethen forecast or measurement.] A cohort of 100 tasks incurs $200 in delivery costs, including failed attempts and verification. Eighty outcomes are independently verified as successful. CPVO is $200 ÷ 80 = $2.50. If the work were priced at $4.50 per verified success, revenue would be 80 × $4.50 = $360, and delivery gross margin (360 − 200) ÷ 360 ≈ 44%, before any costs outside delivery. Now suppose model tokens made up $160 of the $200, and a cheaper model halves token spend to $80, but verified successes fall to 60 and retries and extra verification add $10. Costs become $200 − $80 + $10 = $130, CPVO becomes about $2.17, and revenue falls to $270. Whether that is better depends on what the 20 lost successes were worth to the buyer. CPVO makes that trade visible; token price alone would have called it a clear win.

Cheaper tokens, uncertain outcomes

The illustration generalizes (Figure 3).

Diagram: a cheaper model or lower token price on the left leads to three opposing effects in the middle: more tokens per attempt (longer reasoning, more tool calls), more attempts per task (retries, escalations), and a different verified success rate. These feed into CPVO on the right, which can rise or fall. A note says only end-to-end measurement determines the sign.

Figure 3. Why cheaper tokens need not mean cheaper outcomes. A lower price per token can be offset by longer reasoning, more retries, more verification and lower verified success. CPVO captures the net effect; token price alone does not. The paths shown are mechanisms, not measurements. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis; mechanisms only.

A lower price per token reduces cost only if everything else holds. In practice, cheaper models may reason for longer, call more tools, fail more often and trigger retries or escalations. Each of these raises the cost per attempt or the number of attempts per success. Reasoning-heavy models can be more expensive per token yet cheaper per correct solution on complex problems, which is one of the findings of the cost-of-pass analysis (Erol et al.). Which effect dominates depends on the task, and only end-to-end measurement settles it. Cascades that try a cheap model first can cut cost substantially on some workloads (Chen et al.), but the saving is real only if the cascade's verified success holds up. This is why the decision layer described in Faros optimizes verified outcomes per unit cost rather than tokens, and why the routing literature should report CPVO rather than token savings (Why Learned AI Model Routing Must Beat Good Rules).

Verifier error biases CPVO

The denominator depends on the verifier, and verifiers make mistakes. Let a verifier have false-accept rate f and false-reject rate g. If the true success rate is p, the verified rate is p<sub>obs</sub> = p(1 − g) + (1 − p)f. A lenient verifier with high f inflates the denominator and makes CPVO look lower than it is. A strict verifier with high g does the opposite. When f and g are known from calibration, the true rate can be estimated as (p<sub>obs</sub> − f) ÷ (1 − f − g), provided f + g < 1. That correction requires the error rates, which is one more reason they must be measured: see Evaluating the Evaluators. Comparing CPVO across systems that use different verifiers is meaningless unless the verifiers are calibrated on the same gold set.

Delayed outcomes and censoring

Many outcomes are not final when a task ends. A merge may be reverted, a ticket reopened, a refund disputed. CPVO computed early will count provisional successes that later fail, and CPVO computed late will be stale. Two practices help. First, report CPVO at a stated maturity, for example on tasks whose windows have closed, alongside the share still pending. Second, track how often provisional successes flip, by family. A family where many successes later reverse needs a longer window before its CPVO means anything. Storage of versioned outcomes that makes this possible is described in The Outcome Warehouse.

Where CPVO comes from operationally

CPVO can be computed only if every task's costs and verified outcome are recorded in one place. The Work Receipt is designed for this: its cost fields supply the numerator and its verification fields the denominator, and receipts must reconcile with independent usage meters. Without such a record, CPVO estimates are spreadsheet reconstructions whose errors are unknown.

CPVO and pricing

CPVO is a cost metric, but it naturally suggests pricing per verified outcome. Outcome-based pricing can align a vendor's incentives with a buyer's, and it is used for some categories of automated work. It also creates a conflict of interest: the vendor is paid according to its own verifier, and any proxy that determines payment will attract optimization against it (Skalse et al.; Manheim & Garrabrant). Several safeguards follow. Prefer deterministic verification for billed outcomes, and never let a model judge alone decide what is billed. Let the buyer pin the verifier version. Provide a dispute process that uses the task record as evidence. Audit verifier calibration independently on a sample. Outcome pricing is appropriate only where success is contract-defined, observable and attributable to the work. Broad business outcomes that depend on the customer's own actions are poor candidates. Nothing in this note proposes particular prices.

What CPVO does not capture

CPVO measures the cost of producing successful outcomes. It does not measure their value, which differs between a resolved password reset and a resolved production outage. It does not capture latency, which can matter as much as cost. It does not capture reliability across repeated attempts, which pass^k-style measures do (Yao et al.), or the length of task an agent can handle, which horizon measures track (Kwa et al.). It does not capture risk: two systems with equal CPVO can differ greatly in how often they produce harmful effects. It should be reported alongside quality, latency, intervention rate and incident measures, never as a single number that decides everything.

Limitations

This note defines a metric; it reports no measurements. Allocation of shared overhead is a judgment that affects comparisons. The verifier-bias correction assumes calibrated error rates that hold on the deployed distribution, which may not be true. Ethen's own CPVO is unknown; internal planning documents call measuring it on real or pilot work the most valuable missing evidence. Earlier internal documents used the name cost per successful outcome (CPSO) for a closely related quantity; CPVO makes the verification standard explicit.

Conclusion

The economics of AI agents should be measured in the unit buyers care about: work that was done and can be shown to have been done. CPVO puts every delivery cost, including failures and verification, over independently verified successes, and keeps pending work in view. It is harder to measure than cost per token. That difficulty is the point: it forces the measurement of success that cheaper units let organizations skip.

FAQ

How is CPVO different from cost-of-pass? Cost-of-pass is the expected cost of generating a correct solution, defined for model evaluation. CPVO applies the same idea to deployed agent work, with success defined by independent verification against acceptance criteria and costs that include tools, infrastructure, verification and human review.

Should research costs be included? Not in delivery CPVO. Research and training investment should be reported separately so that delivery margins are not flattered or obscured.

Does lower CPVO always mean a better system? No. CPVO must be read alongside quality, latency, risk and the value of each outcome.

References

  1. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., Narayanan, A. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
  2. Erol, M. H. et al. (2025). Cost-of-Pass: An Economic Framework for Evaluating Language Models. arXiv:2504.13359. https://arxiv.org/abs/2504.13359
  3. Chen, L., Zaharia, M., Zou, J. (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176. https://arxiv.org/abs/2305.05176
  4. Skalse, J. et al. (2022). Defining and Characterizing Reward Hacking. arXiv:2209.13085. https://arxiv.org/abs/2209.13085
  5. Manheim, D., Garrabrant, S. (2018). Categorizing Variants of Goodhart's Law. arXiv:1803.04585. https://arxiv.org/abs/1803.04585
  6. Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
  7. Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Infrastructure

    Reserving a Budget for Verification

    An agent that spends its whole budget doing the work has nothing left to check it. Ethen's mission system reserves verification capacity first — computed in exact integer arithmetic.

  • Company

    Why Ethen Research Lab Publishes Its Work in Public

    Ethen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.

  • Models & Intelligence

    Why Ethen Is Investing in Model Intelligence

    Ethen invests in model intelligence because every decision about which AI model to use — made by a person, by Ethen's AI Gateway, or by Ethen's automatic model choice — is only as good as the facts behind it. Model intelligence is the knowledge needed to make that decision: what a model can do, what it costs, how it performs on which kinds of task, and where each of those facts came from. The model landscape changes too quickly, and public comparisons hide too much, for that knowledge to be assembled ad hoc. So Ethen builds it deliberately: every fact with a source and an owner, Unknown shown instead of guesses, eligibility decided before preference, and a long-term aim of connecting model choices to whether the resulting work actually succeeded.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.