Methods Paper · 2026-10-03 · Evaluation & Verification
Evaluating the Evaluators: Reward Integrity for AI Agents
Every learning signal for an agent passes through an evaluator. If we do not know how often the evaluator is wrong, we do not know what the agent is learning.
Abstract
Agents improve by optimizing against something: a test suite, a rubric, a model judge, a human approval. We call any such mechanism a verifier, and we call the property that its verdicts can safely drive learning reward integrity. This methods paper argues that reward integrity must be measured, verifier by verifier and task class by task class, before any policy or model is trained on verifier output. We show how false accepts and false rejects become learned behavior. We review evidence that proxy rewards are exploitable in principle and in practice, and that model judges carry systematic biases. We then propose six measurements: false-accept rate, false-reject rate, abstention, calibration, drift and exploitability. We also propose a gating rule that blocks reward-based training until those measurements exist. We discuss deterministic versus model graders, expert disagreement and the conflict of interest that arises when an organization is paid on its own verifiers. The paper synthesizes Ethen internal research with published literature. It reports no new measurements.
Why evaluators decide what is learned
An agent cannot observe its true success directly. It observes a verdict. When the verdict comes from a test suite, a reconciliation check, a rubric or a model judge, the agent's learning signal is only as accurate as that mechanism. This is obvious, yet routinely ignored in practice. Benchmarks report agent scores, but rarely the error rates of the graders that produced them. Production systems learn from "accepted" outcomes without measuring how often acceptance was wrong.
The stakes rise as learning moves up the improvement ladder described in Verified Adaptive Intelligence. Using a noisy verifier to rank two prompt variants is a recoverable mistake. Using the same verifier as the reward for reinforcement learning is an instruction to find its blind spots.
From verifier error to learned behavior
Figure 1 shows the mechanism.
Figure 1. How verifier error becomes learned behavior. A learner sees only the verifier's verdict. Where the verifier falsely accepts, optimization pressure pushes the policy toward those regions; where it falsely rejects, good behavior is extinguished. The observed success rate can rise while true success stays flat. Evidence label: CONCEPTUAL DIAGRAM. Source: Ethen internal synthesis; conceptual.
Let a task class have true success rate p. Let a verifier have false-accept rate f (the fraction of failures it accepts) and false-reject rate g (the fraction of successes it rejects). The success rate the learner observes is
p<sub>obs</sub> = p(1 − g) + (1 − p)f.
Two consequences follow. First, when f is non-zero, a policy can increase p<sub>obs</sub> without increasing p by moving toward outputs the verifier mistakenly accepts. Second, f and g are not constants. They vary across regions of the output space, and optimization moves the policy into regions where the verifier was never calibrated. A verifier that is 97% accurate on the distribution it was tested on can be far less accurate on the distribution an optimized policy produces.
What the literature establishes
Reward hacking is structural. Skalse et al. give a formal definition of reward hacking. They show that, outside of narrow conditions, a proxy reward cannot be guaranteed unhackable with respect to the true reward. Gao et al. measured how optimizing against a learned reward model first improves and then degrades performance on a gold-standard reward as optimization pressure increases. They characterize this overoptimization with scaling laws. Pan et al. map how misspecified rewards produce misaligned behavior, including phase transitions where more capable agents exploit misspecification more sharply. The broader phenomenon has a long name in economics. Manheim and Garrabrant categorize the ways a measure stops tracking its target once it becomes the target.
Model judges have systematic biases. LLM judges can agree with humans at rates comparable to inter-human agreement on some tasks, but they show position, verbosity and self-enhancement biases (Zheng et al.). They can be steered by the order in which candidates are presented (Wang et al.). They recognize and favor their own generations (Panickssery et al.). A study of thirteen judge models in a setting where humans agree strongly found that only the largest judges approached reasonable alignment. Even those trailed inter-human agreement, could assign scores differing from human scores by up to five points, and showed a tendency toward leniency (Thakur et al.). Reward models evaluated on dedicated benchmarks show uneven reliability across categories (Lambert et al.).
Deterministic checks are not automatically valid. Tests can be incomplete, assertions can be vacuous, and an end-state check can miss collateral damage. Benchmarks that check for unexpected state changes as well as intended ones exist because naive end-state checks are not enough (Trivedi et al.). A deterministic verifier is easier to audit than a model judge, not exempt from audit.
Abstention helps. Judges that assess their own confidence, decline to judge when uncertain and escalate to stronger judge models can offer provable bounds on agreement with human judgment at a user-specified level (Jung et al.). Statistical methods such as prediction-powered inference combine a small set of human labels with a large set of model predictions to produce valid confidence intervals (Angelopoulos et al.). Retrieval-augmented evaluation frameworks have applied this idea to automatic evaluation (Saad-Falcon et al.).
Six measurements
We propose that every verifier used for learning or billing be characterized by six measurements, per task class (Figure 2).
Figure 2. Six reward-integrity measurements. What a reward-integrity program measures for each verifier, why, and the evidence required before that verifier may supply training rewards. Thresholds are deliberately absent: they depend on the consequences of error for each task class. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen methods proposal.
False-accept rate. Measured on a blind gold set: cases judged by domain experts who do not see the verifier's verdict. The gold set must include near misses, meaning outputs that look correct but are not, because those are where false accepts concentrate. Ethen's internal planning proposes gold sets of 100 to 200 expert-judged items per task class as a starting point [PROPOSED TARGET]. Even that is small. Zero false accepts observed in 100 cases still leaves an upper 95% confidence bound near 3% under independence (Hanley & Lippman-Hand). Small gold sets bound error rates loosely, and reports should say so.
False-reject rate. On the same gold set, stratified by solution strategy. A verifier that accepts only one valid approach suppresses others.
Abstention. The fraction of cases where the verifier declines to judge, and its accuracy when it does judge. Unjudged cases must not become labels by default.
Calibration. Whether the verifier's confidence predicts its correctness. Calibration lets a system set thresholds and route uncertain cases to experts.
Drift. Changes in error rates across verifier versions, model versions and input distributions. Drift is detected by re-running the verifier on a frozen anchor set with every version change.
Exploitability. The rate at which an adversarial agent, explicitly searching for outputs the verifier accepts but experts reject, succeeds. This is the closest measurable proxy for reward-hacking risk. An exploit audit should be repeated whenever the verifier or the policy changes substantially.
The measurement protocol for these quantities is developed in How Should We Measure the Reliability of LLM Verifiers?.
Deterministic and model graders
Ethen's internal architecture work orders verifiers by trust: deterministic invariants, then programmatic verifiers, then expert rubrics, then calibrated model judges. One rule binds them: a model judge never overrides a deterministic policy failure. If a test fails or a policy check denies an action, no judge's opinion converts that into success.
The ordering is not a claim that model judges are useless. They cover open-ended work that deterministic checks cannot reach, such as research synthesis, writing quality and design. The claim is narrower. A model judge's verdicts may supply rewards only after the six measurements exist for the relevant task class, and judges from the same model family as the policy should not grade it. Combining a deterministic skeleton (did the required state change happen?) with a structured judge for semantic dimensions (is the explanation adequate?) often gives better coverage than either alone. Benchmarks for workflow agents increasingly take this approach, using deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions (Li et al., Claw-Eval-Live).
Graders inside benchmarks
Benchmarks are verifiers too. When a benchmark reports that an agent solved 60% of tasks, that number inherits the error of the benchmark's graders, and comparisons between agents inherit it twice. We therefore treat grader characterization as part of benchmark design, not an afterthought. The Ethen VerifiedWork framework requires every track to publish, alongside agent scores, the false-accept and false-reject rates of its graders on a sealed gold set, and to report agent scores with intervals that account for grader error where possible. The same discipline applies to the outcome labels a production system accumulates. In the terms of From AI Traces to Verified Experience, each label should carry its tier and its verifier's version so that later consumers can weight or exclude it.
Expert disagreement
Gold sets are judged by humans, and humans disagree. Disagreement is information, not noise to be averaged away. A program should report inter-rater agreement with an appropriate statistic (Cohen). It should separate disagreement caused by ambiguous task specifications from disagreement among experts about quality. And it should exclude, or analyze separately, items without adequate agreement. If experts cannot agree on whether a task succeeded, no verifier can be calibrated against them, and the task class may not be ready for learning at all.
The reward-integrity gate
These measurements become useful when they gate decisions. We propose the rule in Figure 3.
Figure 3. The reward-integrity gate. Proposed gating rule: a verifier may supply rewards for policy or weight updates only after its error rates are published on a gold set, it has survived an exploit audit, and drift monitoring is in place. Until then its output can inform dashboards but not training. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen methods proposal.
A verifier may supply rewards for policy optimization or weight updates only after three conditions hold for the relevant task class. Its false-accept and false-reject rates are published on a gold set. It has passed an exploit audit. Drift monitoring is in place. Until then, its output may inform dashboards and analysis but not training. Any change to the verifier or a substantial shift in the task distribution triggers re-certification. This is the gate in the improvement ladder of Verified Adaptive Intelligence, and it is also the gate in the decisive learning protocol, How to Test Whether Verified Experience Improves AI Agents.
The conflict of interest in outcome pricing
When an organization is paid per verified outcome, it grades its own work. Unless this is managed, verifier error becomes a billing error that always favors the vendor. Several controls follow from the measurements above: the customer can pin the verifier version used for billing; a dispute channel uses the Work Receipt as evidence; a third party audits verifier calibration periodically on a sample; and deterministic checks are preferred for billed outcomes, with model judges never deciding billing alone. The economics of verified outcomes are developed in Cost Per Verified Outcome.
Why verification gains value as models improve
A natural objection is that better models will make verifiers less necessary. We expect the opposite. More capable agents take longer and more consequential actions, which makes undetected errors more costly. They also become more effective at finding verifier blind spots when optimized against them. And as generation gets cheaper, verification becomes a larger share of the cost of a trustworthy outcome. The argument is developed in Why Better Foundation Models May Make Evaluation More Valuable, Not Less.
Limitations
This paper proposes measurements and a gate; it does not report measured verifier error rates for any Ethen system. Gold sets are expensive and, at practical sizes, bound error rates only loosely. Exploit audits find exploits that their designers think to look for; absence of a found exploit is not absence of exploits. Expert judgment is itself fallible and expensive in specialized domains. And the gate is conservative by design: it will slow learning in domains where the measurement burden is high.
Conclusion
The evaluator is part of the learning system, so its errors become the learner's behavior. Reward integrity treats every verifier as an instrument with error rates to be measured, published and monitored. Verifiers that have not been characterized can inform people, but they should not teach machines.
FAQ
What is reward integrity? The property that a verifier's verdicts are accurate enough, on the distribution a policy will produce, to be used as a learning signal without teaching the policy to exploit the verifier.
Are deterministic tests always trustworthy? No. They are easier to audit than model judges, but tests can be incomplete or vacuous. They need the same false-accept measurement.
Can a model judge ever be used as a reward? Yes, for a task class where its error rates are measured on a gold set, it passes an exploit audit, and its drift is monitored. Not before.
Related research
- How Should We Measure the Reliability of LLM Verifiers? — the measurement protocol for verifiers.
- Ethen VerifiedWork: A Benchmark Framework for AI Systems That Take Action — benchmarks depend on grader quality.
- How to Test Whether Verified Experience Improves AI Agents — learning gated on reward integrity.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI — verified outcomes need trustworthy verifiers.
- From AI Traces to Verified Experience — label tiers in verified experience.
- Why Better Foundation Models May Make Evaluation More Valuable, Not Less — why verification gains value as models improve.
References
- Skalse, J. et al. (2022). Defining and Characterizing Reward Hacking. NeurIPS 2022. arXiv:2209.13085. https://arxiv.org/abs/2209.13085
- Gao, L., Schulman, J., Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760. https://arxiv.org/abs/2210.10760
- Pan, A. et al. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. arXiv:2201.03544. https://arxiv.org/abs/2201.03544
- Manheim, D., Garrabrant, S. (2018). Categorizing Variants of Goodhart's Law. arXiv:1803.04585. https://arxiv.org/abs/1803.04585
- Zheng, L. et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685. https://arxiv.org/abs/2306.05685
- Wang, P. et al. (2023). Large Language Models are not Fair Evaluators. arXiv:2305.17926. https://arxiv.org/abs/2305.17926
- Panickssery, A. et al. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076. https://arxiv.org/abs/2404.13076
- Thakur, A. S. et al. (2024). Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges. arXiv:2406.12624. https://arxiv.org/abs/2406.12624
- Lambert, N. et al. (2024). RewardBench: Evaluating Reward Models for Language Modeling. arXiv:2403.13787. https://arxiv.org/abs/2403.13787
- Trivedi, H. et al. (2024). AppWorld. arXiv:2407.18901. https://arxiv.org/abs/2407.18901
- Jung, J. et al. (2024). Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement. arXiv:2407.18370. https://arxiv.org/abs/2407.18370
- Angelopoulos, A. N. et al. (2023). Prediction-Powered Inference. arXiv:2301.09633. https://arxiv.org/abs/2301.09633
- Saad-Falcon, J. et al. (2023). ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. arXiv:2311.09476. https://arxiv.org/abs/2311.09476
- Li, C. et al. (2026). Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows. arXiv:2604.28139. https://arxiv.org/abs/2604.28139
- Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
- Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1). https://doi.org/10.1177/001316446002000104
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Work Receipts: A Verifiable Record for Autonomous AI Work
A technical report proposing the Work Receipt: one signed record of authority, actions, effects, verification, cost and rights for every unit of autonomous AI work.
- Counterfactual Replay for AI Agents
A research proposal for counterfactual evaluation of AI agents: replaying completed tasks under alternative models, tools, context and recovery strategies.
- Why Better Foundation Models May Make Evaluation More Valuable, Not Less
A position paper stress-testing AI evaluation against 10× better models and 10× cheaper inference, and arguing that verification and assurance gain value.
Explained on the Ethen Blog
- Reserving a Budget for Verification
An agent that spends its whole budget doing the work has nothing left to check it. Ethen's mission system reserves verification capacity first — computed in exact integer arithmetic.
- Why Ethen Shows What It Knows—and What It Doesn't
AI products have a built-in honesty problem: their output sounds equally confident whether it is right, wrong, estimated or invented. Ethen treats that as a design problem, not only a model problem. Across the product, we try to keep four states of knowledge apart — known, estimated, unknown and not checked — and to show each for what it is. When a model fact lacks provenance, Ethen Model Intelligence shows Unknown rather than a guess. When an agent's action times out without confirmation, Ethen's mission system records the outcome as unknown rather than as success or failure. When Ethen Research Lab publishes a proposal, it says the proposal is untested. We do this because people rely on AI outputs more than the outputs deserve when uncertainty is hidden, and because a gap shown honestly is more useful than a gap filled with something plausible.
- What “Done” Should Mean for an AI Agent
For an AI agent, "done" should mean that every requirement of the task has been met and that something other than the agent's own report shows it. Precisely: a task is complete when each required obligation is supported by evidence at the level of checking it needs — a passing test, a reconciled record, a confirmed delivery, an approved review — or has been explicitly waived by the person who owns the task. Three refinements make the definition usable. Keep execution success (the steps ran), task success (the outcome was achieved) and business success (it produced value) apart. Treat success as provisional until it can no longer be reversed. And report partial and unknown outcomes as what they are, instead of rounding them up to done.
Explore this topic
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.