Benchmark Design · 2026-10-03 · Evaluation & Verification
Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action
An agent's final message is the least reliable evidence of what it did. A benchmark for agents that act should grade the state of the world, the record of actions, the cost and the reliability, and should say plainly what its scores cannot prove.
Abstract
Agent benchmarks have moved from question answering toward executable environments: software repositories, web applications, operating systems, simulated businesses. Several problems persist. Many benchmarks freeze their task sets, so they age as models train on them. Many grade final responses rather than verified effects. Grader quality is rarely reported. Cost is often omitted. And reward-design flaws can distort measured performance substantially: one audit found that issues in agentic benchmarks could under- or overestimate performance by up to 100% in relative terms. This paper describes the design of Ethen VerifiedWork, an AI agent benchmark framework for systems that take consequential action. VerifiedWork consists of a core work track and four specialist tracks: Recovery, Transfer, Context and Control. Its principles are: executable environments with real effects; state-based verification ordered from deterministic checks to calibrated judges; cost per verified outcome as a first-class metric; reliability across repeated trials; permission and mandate constraints in every task; contamination controls; and a reporting standard that includes each track's limits. VerifiedWork is a proposed design. It has not yet been run, and no scores are reported.
Why another benchmark framework
Executable agent benchmarks have advanced quickly. SWE-bench grades code changes against repository tests (Jimenez et al.). WebArena provides realistic websites for web agents (Zhou et al.). OSWorld supplies real operating-system environments with execution-based evaluation (Xie et al.). WorkArena measures knowledge-work tasks on an enterprise software platform (Drouin et al.). AppWorld grades with state-based unit tests that also check for unexpected collateral changes (Trivedi et al.). τ-bench evaluates tool–agent–user interaction against database state and introduced pass^k to measure reliability across repeated trials (Yao et al.). Its successor models settings where both agent and user can act on shared state (Barres et al.). TheAgentCompany simulates a small software company (Xu et al.), and live benchmarks refresh tasks from current workflow demand (Li et al., Claw-Eval-Live).
Three gaps remain for organizations that deploy agents on consequential work.
Benchmark validity is under-examined. An audit of agentic benchmarks found problems in task setup and reward design. Examples include insufficient test cases in one widely used coding benchmark, and an evaluation that counted empty responses as successful. Such issues were estimated to under- or overestimate performance by up to 100% in relative terms [EXTERNAL PRIMARY-SOURCE RESULT] (Zhu et al., Agentic Benchmark Checklist). A deterministic grader is not automatically a valid one.
Cost and reliability are secondary. Accuracy-focused benchmarking encourages needlessly complex and costly agents and can lead to mistaken conclusions about where gains come from (Kapoor et al.). A system that succeeds once in three attempts at high cost is not production-ready, however good its best run looks.
Authority is absent. Few benchmarks test whether an agent stays within delegated authority, honors approvals, respects revocation and handles unknown effects safely. In enterprise deployment, these failures matter at least as much as task failure.
VerifiedWork is designed around these gaps. It complements existing benchmarks rather than replacing them, and it can incorporate their environments where licenses allow.
Design principles
1. Grade effects, not narratives. Every task's success criteria are defined over the state of the environment after the agent acts, and over the record of actions it took. The agent's final message is evidence of what it claims, which is useful for measuring false completion (see Commitment Graphs). It is never evidence of success.
2. Order verification by trust. Graders apply deterministic state checks first, then programmatic checks, then calibrated model judges, and the judges are used only for semantic dimensions that deterministic checks cannot reach. A model judge never overrides a failed deterministic check. Every grader's false-accept and false-reject rates on a sealed gold set are published. The certification procedure is How Should We Measure the Reliability of LLM Verifiers?, and the broader rationale is Evaluating the Evaluators.
3. Report cost per verified outcome. Every result includes the full cost of producing verified successes, including failed attempts and the cost of any verification the agent itself performs, as defined in Cost Per Verified Outcome.
4. Measure reliability. Each task is run several times; results report pass^k alongside mean success.
5. Constrain authority in every task. Each task runs under an explicit mandate: permitted tools and resources, a budget, and approval requirements. Violations are scored even when the task otherwise succeeds.
6. Contain and observe effects. Environments are executable and stateful, but every side effect is contained within the environment, and a Work Receipt records every action.
7. Pin everything. Results are reproducible only if the environment, task set, grader, model, tools and harness are versioned and reported.
Task specification
Each VerifiedWork task specifies:
- the initial state of every system in the environment;
- the allowed tools and permissions, expressed as a mandate;
- the budget: money, tokens, time and allowed human interventions;
- success and failure criteria over final state and action record, including forbidden collateral changes;
- expert difficulty and task-family lineage, so that splits can be made by family;
- dependency versions for every service and tool;
- permitted network behavior;
- the grader version and its calibration record;
- the rights under which the task and its data may be used.
Figure 1 shows how a task is run and scored.
Figure 1. How a VerifiedWork task is run and scored. A task specification fixes the initial state, permissions, budget and success criteria. The agent acts in an executable environment; graders read the resulting state, the action record and the cost, never only the agent's final message. Evidence label: EXPERIMENT DESIGN. Source: Ethen benchmark design (proposed).
Tracks
VerifiedWork is one suite with a core track and four specialist tracks, each with its own design paper (Figure 2).
Figure 2. Tracks, what each measures, and what it cannot prove. VerifiedWork is one suite with a core track and four specialist tracks, each described in its own design paper. The right-hand column is part of the design: every track publishes its limits alongside its scores. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen benchmark design (proposed).
Core work measures verified completion of multi-step tasks with real effects in code, operational workflows and evidence-based research.
Recovery injects failures, including tool errors, stale state, ambiguous effects after timeouts and partial completion, and measures whether agents recover, reconcile, escalate or stop safely.
Transfer measures whether capabilities, skills and configurations keep working when model or tool versions change.
Context measures whether agents retain obligations, permissions and evidence across long tasks and context compaction.
Control measures delegation, approval binding, revocation and budget enforcement.
Each track publishes what its scores cannot prove. This is part of the design, not a disclaimer. A Control score does not certify compliance. A Transfer score does not predict behavior on models that were not tested. A Core score does not establish product-market fit or general intelligence.
Environments
VerifiedWork needs environments that are executable, resettable, rich enough to support consequential tasks, and faithful enough that results mean something. Three families are planned:
- Code environments: containerized repositories with tests, build systems and continuous-integration checks. Verification is strongest here.
- Enterprise workflow environments: a simulated organization with ticketing, customer records, documents, approvals, identity and permissions. The design is described in Ethen Synthetic Enterprise.
- Research-evidence environments: frozen source collections where claims must be supported by specific passages, graded by citation checks and expert rubrics.
Environment quality matters for training as well as evaluation. Reinforcement learning in a high-fidelity enterprise simulation has been reported to produce gains that transfer to out-of-distribution benchmarks, with the authors attributing transfer to task-centric world building, expert rubrics and realistic workflows (Mehta et al.). Fidelity must be tested, not assumed. Each environment carries probes that compare its behavior with the real systems it imitates.
Scoring rules
How per-task results become summary scores determines what a benchmark rewards, so the rules are fixed in advance.
Violations gate success. A task in which the agent commits a critical violation, such as acting outside its mandate, executing an unapproved irreversible action or duplicating an effect, scores as a failure regardless of whether the end state matches. Violations are also reported separately, by severity class, so that a system with high success and occasional critical violations cannot hide them in an average.
Partial credit is separate. Many tasks have several obligations. The share of obligations met is reported as its own metric, alongside, never instead of, binary verified success. This keeps "almost done" visible without letting it inflate the headline.
Families are weighted equally. Summary scores average across task families with equal weight, so that a benchmark dominated by one easy family does not dominate the score. Scores weighted to a declared traffic mix may be reported as well, labeled as such.
Intervals are clustered. Confidence intervals are computed by resampling at the family level, because tasks within a family share structure and are not independent.
Cost is normalized but not hidden. Cost per verified outcome is reported at fixed, dated prices, so that results remain comparable when providers change their prices.
Required baselines
A score means little without baselines. Every VerifiedWork release reports at least four:
- a null agent that takes no actions, which catches tasks that pass without work, the class of grading flaw the benchmark-validity audit found;
- a scripted baseline, where a deterministic script can complete part of a family, to show how much of the family requires judgment;
- a direct frontier-model baseline, the same model called with the same tools but without the system's surrounding harness, routing (the decision layer described in Faros) or skills, so that the system's contribution can be separated from the model's;
- where feasible, a human expert reference on a sample of tasks, with time and cost recorded.
The direct-model baseline matters most for any organization that builds systems around models. If a system cannot beat the model it wraps on verified outcomes per unit cost, its additions are not earning their keep.
Splits and contamination
Benchmarks lose validity when models are trained on them. Contamination of static benchmarks is well documented: a carefully matched replacement for a grade-school math benchmark revealed accuracy drops of up to 8% for some models (Zhang et al.). Live benchmarks that refresh their tasks address this directly (White et al.; Jain et al.). VerifiedWork uses:
- family and temporal splits, so that held-out tasks differ in kind and in time from development tasks;
- a public slice for reproducibility and a sealed set for decisive results, managed by a reviewer independent of any system being evaluated;
- canary strings in every task and near-duplicate exclusion against training corpora where Ethen controls them;
- rolling refresh, with compromised tasks retired;
- a frozen anchor suite for longitudinal comparison alongside an evolving live suite.
Reporting standard
A VerifiedWork result is a scorecard, not a number (Figure 3).
Figure 3. Reporting standard for every result. A VerifiedWork result is not a single number. Each submission reports these fields, and results lacking them are not comparable. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen benchmark design (proposed).
Every submission reports verified success with intervals by task family; reliability across trials; cost per verified outcome; latency; policy and permission violations; recovery outcomes and human interventions; grader versions with published error rates; exact versions of every component; and a contamination disclosure. Results that omit fields are not comparable and are not ranked.
Governance
Benchmarks run by an organization that also builds agents invite skepticism, and they should. Four rules apply. First, no self-review: people who build a system under test do not grade its decisive results. Second, publish losses: when Ethen-built systems perform worse than alternatives, the results are published with the same prominence. Third, open methodology: task specifications for the public slice, grader definitions and scoring code are released. Fourth, independent audit: a sample of sealed results is re-graded by an external party.
Limitations
VerifiedWork has not been run. Environment fidelity, grader calibration and task coverage are open engineering problems, and the first versions will be narrow. Simulated enterprise environments may not reflect the complexity of real organizations. Sealed sets limit outside reproducibility, which the public slice only partly offsets. A benchmark designed by an organization that also builds agents carries a conflict of interest that governance rules mitigate but do not remove.
Conclusion
Agents that act should be evaluated on what they changed, at what cost, under what authority and how reliably, by graders whose errors are known, on tasks they have not seen. VerifiedWork is a design for doing that across five tracks, with each track's limits published alongside its scores. Whether it becomes useful depends on execution: faithful environments, calibrated graders and the discipline to publish uncomfortable results.
FAQ
How is VerifiedWork different from existing agent benchmarks? It combines state-based grading with cost per verified outcome, reliability across trials, explicit authority constraints, recovery and transfer tracks, published grader error rates, and a "what this cannot prove" statement for every track.
Will the tasks be public? A public slice will be released for reproducibility. Decisive results use a sealed set managed independently, to limit contamination.
Can external agents be evaluated? That is the intention. The framework is designed to evaluate any agent system that can act through the environment's tools.
Related research
- VerifiedWork Recovery: Evaluating AI Agents Under Failure and Partial Effects — Recovery track.
- VerifiedWork Transfer: Measuring Capability Across Model and Tool Changes — Transfer track.
- VerifiedWork Context: Measuring What AI Agents Must Remember — Context track.
- VerifiedWork Control: Evaluating Delegation, Approval, Revocation, and Agent Authority — Control track.
- Ethen Synthetic Enterprise: An Executable World for Enterprise-Agent Research — environment substrate.
- Evaluating the Evaluators: Reward Integrity for AI Agents — grader integrity.
- How Should We Measure the Reliability of LLM Verifiers? — how graders are certified.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI — cost reporting.
References
- Zhu, Y. et al. (2025). Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825. https://arxiv.org/abs/2507.02825
- Jimenez, C. E. et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770. https://arxiv.org/abs/2310.06770
- Zhou, S. et al. (2023). WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv:2307.13854. https://arxiv.org/abs/2307.13854
- Xie, T. et al. (2024). OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv:2404.07972. https://arxiv.org/abs/2404.07972
- Drouin, A. et al. (2024). WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? arXiv:2403.07718. https://arxiv.org/abs/2403.07718
- Trivedi, H. et al. (2024). AppWorld. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
- Yao, S. et al. (2024). τ-bench. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Barres, V. et al. (2025). τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982. https://arxiv.org/abs/2506.07982
- Xu, F. F. et al. (2024). TheAgentCompany. arXiv:2412.14161. https://arxiv.org/abs/2412.14161
- Li, C. et al. (2026). Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows. arXiv:2604.28139. https://arxiv.org/abs/2604.28139
- Kapoor, S. et al. (2024). AI Agents That Matter. arXiv:2407.01502. https://arxiv.org/abs/2407.01502
- Mehta, S. et al. (2026). EnterpriseBench CoreCraft. arXiv:2602.16179. https://arxiv.org/abs/2602.16179
- Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332
- White, C. et al. (2024). LiveBench. arXiv:2406.19314. https://arxiv.org/abs/2406.19314
- Jain, N. et al. (2024). LiveCodeBench. arXiv:2403.07974. https://arxiv.org/abs/2403.07974
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Work Receipts: A Verifiable Record for Autonomous AI Work
A technical report proposing the Work Receipt: one signed record of authority, actions, effects, verification, cost and rights for every unit of autonomous AI work.
- Evaluating the Evaluators: Reward Integrity for AI Agents
A methods paper on reward integrity for AI agents: verifier false accepts and rejects, abstention, grader drift, expert disagreement and reward hacking.
- Counterfactual Replay for AI Agents
A research proposal for counterfactual evaluation of AI agents: replaying completed tasks under alternative models, tools, context and recovery strategies.
Explained on the Ethen Blog
- Making Mission Completion Depend on Evidence
In the Stage-0 mission system, running the work and proving the work are different transitions — and only independent evidence unlocks the second.
- What Makes an AI Agent Job Verifiable
An agent job is verifiable when someone other than the agent can confirm what was supposed to happen, what actually happened, and what remains unknown. This guide turns that idea into a checklist you can apply to any system.
- Why Ethen Is a Family of Specialized AI Apps, Not One App
Ethen is organized as a family of focused apps rather than one universal chat window because different kinds of AI work need different things to persist, different controls, and different evidence. A quick question needs a thread. A research investigation needs sources and branches that survive across sessions. A code change needs a repository, tests and review. A goal-driven job needs a plan, approvals and a record of what actually happened. Ethen gives each kind of work one canonical home — Chat, Code, Studio, Research, Designer, Founder, the Platform products and Desktop — and keeps the account, model access, permissions and evidence shared underneath, so it still works as one Ethen.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.