Research Note · 2026-10-03 · Enterprise / Sovereign AI
Cost Per Verified Outcome: A Better Economic Unit for Agentic AI
Buyers of AI agents do not want tokens. They want work done, and they want to know it was done. The unit of account should say so.
Abstract
AI economics is still mostly measured in units inherited from earlier technology: cost per token, per request or per seat. For agents that perform multi-step work, these units can move in the opposite direction from the value delivered. A cheaper model can use more tokens, trigger more retries and succeed less often, so that work gets more expensive while tokens get cheaper. This note defines cost per verified outcome (CPVO): all eligible delivery costs for a defined cohort and period, divided by the number of outcomes independently verified as successful. We specify what belongs in the numerator, including failed attempts and verification itself. We specify what counts in the denominator: verified successes at the task level, with pending work reported separately. We compare CPVO with other units, relate it to recent proposals for cost-aware evaluation, and work through an illustrative calculation. We also show how verifier error biases the measured value and why outcome-based pricing built on CPVO needs safeguards against the seller grading its own work. CPVO is an Ethen research proposal. Ethen's own CPVO is unknown and has not been measured.
Why the familiar units mislead
Cost per token measures the price of a raw input. It ignores how many tokens a task needs, how often attempts fail, and whether the output was useful. Prices per token have fallen quickly, and published analyses track how the cost of reaching a given capability level has declined. For an organization deploying agents, a falling token price is necessary but not sufficient for falling costs of useful work. If inference becomes ten times cheaper, verification and human review become a larger share of the cost of a trustworthy outcome, an argument developed in Why Better Foundation Models May Make Evaluation More Valuable, Not Less.
Cost per request is closer to the work but still counts failures as if they were successes, and treats a request that resolves a problem in one call the same as one that starts a ten-step chain.
Cost per seat decouples price from cost entirely. It suits tools that augment people, and fits poorly when agents do work autonomously and usage varies by orders of magnitude between customers.
Cost per task counts attempted tasks. It includes failures in the numerator, but rewards attempting tasks rather than completing them.
Cost per successful task is closer still, but only as good as the definition of "successful". If success is self-reported by the agent, or inferred from "the run completed", the metric invites optimism, and eventually gaming.
Figure 1 compares these units on the properties that matter.
Figure 1. Economic units compared. Each unit answers a different question. Only units with a verified-success denominator track what a buyer of agent work receives, and only CPVO makes the verification standard explicit. Marks are qualitative judgments. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative judgment.
Evaluation research has reached a similar conclusion from a different direction. A critique of agent benchmarks argued that a narrow focus on accuracy without attention to cost has produced needlessly complex and costly agents, and has led to mistaken conclusions about the sources of accuracy gains. It recommends optimizing accuracy and cost jointly (Kapoor et al.). An economic framework for language models defines cost-of-pass, the expected monetary cost of generating a correct solution, and tracks how the cheapest available cost-of-pass for different task types has changed over time (Erol et al.). CPVO is the counterpart of these ideas for deployed agent work, where "correct" must be established by verification against acceptance criteria, and where costs include far more than inference.
Definition
CPVO for a cohort and period = (all eligible delivery costs for that cohort and period) ÷ (number of outcomes in that cohort independently verified as successful)
Figure 2 shows the components.
Figure 2. What goes into the numerator and the denominator. The numerator includes every eligible delivery cost for a cohort, including failed attempts and verification. The denominator counts only independently verified successes at the task-success level, with pending work reported separately rather than dropped or counted. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis (CPVO definition).
The numerator includes everything required to deliver the work. That means provider tokens with the effects of caching (cache reads, writes and misses priced as billed), tool and external API calls, sandbox and infrastructure compute, storage and egress, the cost of verification itself (running tests, model judges, expert review), required human review and delivery support, and an allocation of operational overhead. Costs of failed attempts are included, because failures are part of what it costs to produce successes. Research and model-training investment is excluded from delivery CPVO and reported separately. Hiding delivery costs in a research budget makes margins look better than they are.
The denominator counts only verified successes. A success is an outcome that met its acceptance criteria according to a named verifier, at the task-success level, after any reopen or revert window has closed. Execution success ("the calls returned") does not count. Business success ("revenue increased") is reported separately, if measurable at all, because attribution to individual tasks is weak. Three accounting rules keep the denominator honest:
- Pending work is visible. Tasks still inside their verification window are reported as pending, not dropped and not counted. CPVO computed only on settled tasks should state how many were pending.
- Zero is undefined, not zero. A cohort with no verified successes has undefined CPVO, not a CPVO of zero.
- The verifier is named. Because the denominator depends on what counts as verified, every CPVO figure states the verification standard used.
CPVO should be reported with its distribution, such as median and 95th-percentile cost per task by task family, and with interval estimates. Averages across heterogeneous work hide the expensive tail where most cost risk lives.
An illustrative calculation
[ILLUSTRATIVE EXAMPLE — arithmetic only; not an Ethen forecast or measurement.] A cohort of 100 tasks incurs $200 in delivery costs, including failed attempts and verification. Eighty outcomes are independently verified as successful. CPVO is $200 ÷ 80 = $2.50. If the work were priced at $4.50 per verified success, revenue would be 80 × $4.50 = $360, and delivery gross margin (360 − 200) ÷ 360 ≈ 44%, before any costs outside delivery. Now suppose model tokens made up $160 of the $200, and a cheaper model halves token spend to $80, but verified successes fall to 60 and retries and extra verification add $10. Costs become $200 − $80 + $10 = $130, CPVO becomes about $2.17, and revenue falls to $270. Whether that is better depends on what the 20 lost successes were worth to the buyer. CPVO makes that trade visible; token price alone would have called it a clear win.
Cheaper tokens, uncertain outcomes
The illustration generalizes (Figure 3).
Figure 3. Why cheaper tokens need not mean cheaper outcomes. A lower price per token can be offset by longer reasoning, more retries, more verification and lower verified success. CPVO captures the net effect; token price alone does not. The paths shown are mechanisms, not measurements. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis; mechanisms only.
A lower price per token reduces cost only if everything else holds. In practice, cheaper models may reason for longer, call more tools, fail more often and trigger retries or escalations. Each of these raises the cost per attempt or the number of attempts per success. Reasoning-heavy models can be more expensive per token yet cheaper per correct solution on complex problems, which is one of the findings of the cost-of-pass analysis (Erol et al.). Which effect dominates depends on the task, and only end-to-end measurement settles it. Cascades that try a cheap model first can cut cost substantially on some workloads (Chen et al.), but the saving is real only if the cascade's verified success holds up. This is why the decision layer described in Faros optimizes verified outcomes per unit cost rather than tokens, and why the routing literature should report CPVO rather than token savings (Why Learned AI Model Routing Must Beat Good Rules).
Verifier error biases CPVO
The denominator depends on the verifier, and verifiers make mistakes. Let a verifier have false-accept rate f and false-reject rate g. If the true success rate is p, the verified rate is p<sub>obs</sub> = p(1 − g) + (1 − p)f. A lenient verifier with high f inflates the denominator and makes CPVO look lower than it is. A strict verifier with high g does the opposite. When f and g are known from calibration, the true rate can be estimated as (p<sub>obs</sub> − f) ÷ (1 − f − g), provided f + g < 1. That correction requires the error rates, which is one more reason they must be measured: see Evaluating the Evaluators. Comparing CPVO across systems that use different verifiers is meaningless unless the verifiers are calibrated on the same gold set.
Delayed outcomes and censoring
Many outcomes are not final when a task ends. A merge may be reverted, a ticket reopened, a refund disputed. CPVO computed early will count provisional successes that later fail, and CPVO computed late will be stale. Two practices help. First, report CPVO at a stated maturity, for example on tasks whose windows have closed, alongside the share still pending. Second, track how often provisional successes flip, by family. A family where many successes later reverse needs a longer window before its CPVO means anything. Storage of versioned outcomes that makes this possible is described in The Outcome Warehouse.
Where CPVO comes from operationally
CPVO can be computed only if every task's costs and verified outcome are recorded in one place. The Work Receipt is designed for this: its cost fields supply the numerator and its verification fields the denominator, and receipts must reconcile with independent usage meters. Without such a record, CPVO estimates are spreadsheet reconstructions whose errors are unknown.
CPVO and pricing
CPVO is a cost metric, but it naturally suggests pricing per verified outcome. Outcome-based pricing can align a vendor's incentives with a buyer's, and it is used for some categories of automated work. It also creates a conflict of interest: the vendor is paid according to its own verifier, and any proxy that determines payment will attract optimization against it (Skalse et al.; Manheim & Garrabrant). Several safeguards follow. Prefer deterministic verification for billed outcomes, and never let a model judge alone decide what is billed. Let the buyer pin the verifier version. Provide a dispute process that uses the task record as evidence. Audit verifier calibration independently on a sample. Outcome pricing is appropriate only where success is contract-defined, observable and attributable to the work. Broad business outcomes that depend on the customer's own actions are poor candidates. Nothing in this note proposes particular prices.
What CPVO does not capture
CPVO measures the cost of producing successful outcomes. It does not measure their value, which differs between a resolved password reset and a resolved production outage. It does not capture latency, which can matter as much as cost. It does not capture reliability across repeated attempts, which pass^k-style measures do (Yao et al.), or the length of task an agent can handle, which horizon measures track (Kwa et al.). It does not capture risk: two systems with equal CPVO can differ greatly in how often they produce harmful effects. It should be reported alongside quality, latency, intervention rate and incident measures, never as a single number that decides everything.
Limitations
This note defines a metric; it reports no measurements. Allocation of shared overhead is a judgment that affects comparisons. The verifier-bias correction assumes calibrated error rates that hold on the deployed distribution, which may not be true. Ethen's own CPVO is unknown; internal planning documents call measuring it on real or pilot work the most valuable missing evidence. Earlier internal documents used the name cost per successful outcome (CPSO) for a closely related quantity; CPVO makes the verification standard explicit.
Conclusion
The economics of AI agents should be measured in the unit buyers care about: work that was done and can be shown to have been done. CPVO puts every delivery cost, including failures and verification, over independently verified successes, and keeps pending work in view. It is harder to measure than cost per token. That difficulty is the point: it forces the measurement of success that cheaper units let organizations skip.
FAQ
How is CPVO different from cost-of-pass? Cost-of-pass is the expected cost of generating a correct solution, defined for model evaluation. CPVO applies the same idea to deployed agent work, with success defined by independent verification against acceptance criteria and costs that include tools, infrastructure, verification and human review.
Should research costs be included? Not in delivery CPVO. Research and training investment should be reported separately so that delivery margins are not flattered or obscured.
Does lower CPVO always mean a better system? No. CPVO must be read alongside quality, latency, risk and the value of each outcome.
Related research
- Work Receipts: A Verifiable Record for Autonomous AI Work — receipts supply numerator and denominator.
- Evaluating the Evaluators: Reward Integrity for AI Agents — verifier error moves the denominator.
- Faros: Researching How Intelligence Should Choose Intelligence — CPVO as Faros objective.
- Why Better Foundation Models May Make Evaluation More Valuable, Not Less — CPVO when inference is 10× cheaper.
- The Outcome Warehouse: Turning Completed AI Work Into Research Assets — delayed outcomes and censoring.
References
- Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., Narayanan, A. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
- Erol, M. H. et al. (2025). Cost-of-Pass: An Economic Framework for Evaluating Language Models. arXiv:2504.13359. https://arxiv.org/abs/2504.13359
- Chen, L., Zaharia, M., Zou, J. (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176. https://arxiv.org/abs/2305.05176
- Skalse, J. et al. (2022). Defining and Characterizing Reward Hacking. arXiv:2209.13085. https://arxiv.org/abs/2209.13085
- Manheim, D., Garrabrant, S. (2018). Categorizing Variants of Goodhart's Law. arXiv:1803.04585. https://arxiv.org/abs/1803.04585
- Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Ethen Synthetic Enterprise: An Executable World for Enterprise-Agent Research
A research proposal for enterprise agent simulation: an executable synthetic company with CRM, support, documents, identity, approvals, finance and email.
- Process Memory: Learning How Organizations Actually Get Work Done
A research proposal for process memory for AI agents: mining completed, verified work into per-organization process models that guide plans and flag anomalies.
- Tenant Replay: Private Evaluation Inside Enterprise Boundaries
A research proposal for private AI evaluation: replaying an enterprise's own historical agent tasks inside its boundary, with only bounded aggregates leaving.
Explained on the Ethen Blog
- Reserving a Budget for Verification
An agent that spends its whole budget doing the work has nothing left to check it. Ethen's mission system reserves verification capacity first — computed in exact integer arithmetic.
- Why Ethen Research Lab Publishes Its Work in Public
Ethen Research Lab publishes its work in public so that our claims can be checked, not just read. Every publication states what kind of evidence it contains — a measured result, a research synthesis, a proposal, a protocol or a benchmark design — and the first library of 40 papers says plainly that none of them reports a new measured Ethen result. Publishing that way does four things: it holds our claims to the evidence we actually have, lets others inspect our methods, commits us to how a hypothesis will be tested before any data arrive, and keeps research clearly separate from product claims. We also say what we keep private and why.
- Why Ethen Is Investing in Model Intelligence
Ethen invests in model intelligence because every decision about which AI model to use — made by a person, by Ethen's AI Gateway, or by Ethen's automatic model choice — is only as good as the facts behind it. Model intelligence is the knowledge needed to make that decision: what a model can do, what it costs, how it performs on which kinds of task, and where each of those facts came from. The model landscape changes too quickly, and public comparisons hide too much, for that knowledge to be assembled ad hoc. So Ethen builds it deliberately: every fact with a source and an owner, Unknown shown instead of guesses, eligibility decided before preference, and a long-term aim of connecting model choices to whether the resulting work actually succeeded.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.