Skip to content

EthenEthenEthen

Research Protocol · 2026-10-03 · Evaluation & Verification

How Should We Measure the Reliability of LLM Verifiers?

Publication type
Research Protocol
Research program
Evaluation & Verification
Published
Authors
Ethen Research Lab
Reading time
12 min read

Before a verifier's verdicts drive learning, billing or headline results, its error rates must be known: per task class, on the outputs it will actually see, and under attack. This protocol says how.

Cover image for "How Should We Measure the Reliability of LLM Verifiers?". Decorative abstract motif; contains no data.

Abstract

Verifiers decide what counts as success. They include test suites, programmatic checks, rubric graders and, increasingly, language models acting as judges. Their errors become errors in benchmark scores, training rewards and bills. This protocol specifies how to measure LLM verifier reliability for a defined task class. Steps include constructing a stratified gold set that deliberately includes near misses, unsupported but plausible outputs and samples from optimized agents. Experts label the gold set blind, with agreement measured. The verifier is run in standard and perturbed configurations. False-accept and false-reject rates, abstention, calibration and agreement are estimated with intervals. An adversarial exploit audit follows, and the results are recorded in a grader card that grants specific uses: dashboards, evaluation, billing or training reward. We discuss sample sizes for rare errors, distribution shift between gold sets and deployed outputs, and threats to validity. Ethen has not yet measured verifier error rates for its own systems. The working title "How Reliable Are LLM Verifiers?" was replaced with this protocol title because no results exist.

Why a protocol is needed

The case for measuring verifiers before trusting them is made in Evaluating the Evaluators. In brief: proxy rewards can be exploited in principle (Skalse et al.) and in practice (Pan et al.; Gao et al.). LLM judges show position, verbosity and self-preference biases (Zheng et al.; Wang et al.; Panickssery et al.), and even the strongest judges evaluated in one study trailed inter-human agreement and tended toward leniency (Thakur et al.). Deterministic benchmark graders also fail: an audit of agentic benchmarks found insufficient tests and grading flaws that could distort measured performance by up to 100% in relative terms (Zhu et al.).

What is missing is a standard procedure that turns these concerns into numbers for a particular verifier and task class. This protocol is that procedure. It is designed to be run before a verifier is used for anything consequential, and again whenever the verifier, the task distribution or the agents being judged change substantially.

Objectives

For a verifier V and a task class C:

  1. Estimate V's false-accept rate (the share of failures it accepts) and false-reject rate (the share of successes it rejects) on C, with intervals.
  2. Estimate V's abstention coverage and its accuracy when it does judge.
  3. Assess calibration of V's confidence, if it reports one.
  4. Measure V's susceptibility to known biases and to adversarial exploitation.
  5. Decide which uses V is eligible for on C.

Step 1: build a stratified gold set

Figure 1 lists the strata.

Table of seven gold-set strata: clear successes, clear failures, near misses, plausible but unsupported outputs, partial effects, adversarial exploits, and on-policy samples from optimized agents. Columns: what the stratum contains and which error it targets. Near misses and plausible-unsupported outputs target false accepts; clear successes with unusual approaches target false rejects; on-policy samples target distribution shift.

Figure 1. Gold-set strata and why each is included. A gold set drawn only from easy cases flatters every verifier. Stratifying by case type, and including outputs from optimized policies, concentrates measurement where verifier errors occur. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Errors concentrate in hard cases, so a gold set sampled at random from typical outputs will underestimate them. The gold set therefore deliberately includes near misses: work that is almost correct but has a decisive flaw. It includes plausible but unsupported outputs, which are confident and well-formed but not grounded in evidence. It includes partial effects, where some obligations are met and others are not, and adversarial exploits crafted to fool the verifier. It also includes clear successes that use unusual but valid approaches, which reveal false rejects.

Most importantly, it includes on-policy samples: outputs from agents that have been tuned or selected using this verifier or a similar one. A verifier's error rate on outputs from an un-optimized agent says little about its error rate on outputs from an agent that has learned where the verifier is lenient. Gold sets should be refreshed with new on-policy samples whenever the agents being judged change.

Ethen's internal planning proposes 100 to 200 expert-labeled items per task class as a starting point [PROPOSED TARGET]. The next step explains why that is a floor, not a ceiling, for safety-relevant classes.

Step 2: label blind, measure agreement

Each item is labeled by at least two domain experts who see the task, the output and the environment state, but not the verifier's verdict. Labels distinguish execution success, task success and, where relevant, business success, following the success levels used across this library. Disagreements are adjudicated by a third expert. Inter-rater agreement is reported with Cohen's κ for categorical labels (Cohen).

Items on which experts cannot agree after adjudication are set aside and reported separately. They are not forced into the gold set. A high rate of irreducible disagreement means the task class's success criteria are ambiguous, and no verifier can be calibrated against them until the criteria are clarified.

Step 3: run the verifier, standard and perturbed

The verifier is run on every gold item in its standard configuration. To measure known biases, it is also run in perturbed configurations: candidate order swapped for pairwise judges, output length varied while content is held fixed, paraphrased outputs, and, for model judges, judge models from the same family as the generating agent versus different families. Differences between standard and perturbed verdicts are reported as bias measures. Results are reported per category within the task class, because reward models and judges show uneven reliability across categories (Lambert et al.).

Step 4: estimate error rates with intervals

False-accept and false-reject rates are estimated per task class with binomial intervals, accounting for clustering where gold items share tasks or sources. Two features of the estimates deserve emphasis.

Rare errors need large samples. If a verifier accepts none of the n failures in a gold set, its true false-accept rate may still be as high as about 3/n at 95% confidence (Hanley & Lippman-Hand). [ILLUSTRATIVE EXAMPLE — arithmetic only.] With 60 failures in a 150-item gold set, zero observed false accepts is consistent with a true rate near 5%. For task classes where a false accept means a billed failure or a harmful action, that bound may be too loose, and the gold set must grow.

Downstream metrics inherit verifier error. When a verifier is used to score many outputs, prediction-powered inference combines its verdicts with a smaller set of expert labels to produce valid confidence intervals for the downstream metric (Angelopoulos et al.). Automated evaluation frameworks have applied the same idea (Saad-Falcon et al.). Where gold labels exist for a sample of production outputs, downstream success rates should be reported with such intervals rather than as raw verifier pass rates.

Abstention. For verifiers that can decline to judge, coverage and accuracy-when-judging are reported together. A verifier that judges only easy cases looks accurate. Selective evaluation methods can offer provable bounds on agreement with humans at a chosen level, escalating uncertain cases to stronger judges (Jung et al.). The protocol measures whether those bounds hold on the task class in question.

Step 5: exploit audit

An adversarial agent is given the verifier's interface, a budget and the goal of producing outputs that the verifier accepts but experts reject. Its success rate, the budget it needed, and representative exploits are recorded. A verifier that withstands a modest attacker has passed something. One that falls quickly is unsuitable as a training reward regardless of its gold-set accuracy, because optimization will find what the attacker found. The audit is repeated when the verifier changes or when the agents it judges become substantially more capable.

Step 6: the grader card and eligibility

Figure 2 shows the measurement pipeline, and Figure 3 the contents of the resulting grader card.

Pipeline: case sourcing across strata; blind expert labeling by two or more raters with adjudication; agreement check, with low-agreement items set aside; verifier runs in standard and perturbed configurations (order swap, length control, paraphrase); metrics computed per task class (false accept, false reject, abstention, calibration, agreement); exploit audit by an adversarial agent; grader card; eligibility decision for dashboard, evaluation, billing or training reward.

Figure 2. Measurement pipeline. Experts label the gold set blind to the verifier. The verifier is run in standard and perturbed configurations, including position swaps and length controls. Results produce a grader card whose contents determine what the verifier may be used for. Evidence label: EXPERIMENT DESIGN. Source: Ethen research protocol (proposed).

Template table for a grader card with nine fields: verifier identity and version; task classes covered; gold-set size and strata per class; false-accept rate with interval; false-reject rate with interval; abstention coverage and accuracy when judging; calibration summary; exploit-audit result and date; eligible uses per task class.

Figure 3. Grader card: required contents. Every certified verifier ships with a card. Uses are granted per task class; a verifier can be eligible for dashboards in one class and for training rewards in another. Evidence label: PROPOSED MEASUREMENT FRAMEWORK. Source: Ethen research protocol (proposed).

Eligibility is granted per task class, in increasing order of stringency:

  • Dashboard and analysis: any verifier with a published card.
  • Evaluation results reported externally: error rates with intervals published alongside the results.
  • Billing: deterministic or programmatic verification preferred, with model judges never deciding billing alone. Error-rate bounds must be acceptable to the customer, who may pin the verifier version.
  • Training reward: published error rates, a passed exploit audit and drift monitoring in place. This is the reward-integrity gate described in How to Test Whether Verified Experience Improves AI Agents.

A worked example

[ILLUSTRATIVE EXAMPLE — a design sketch, not a measurement.] A model judge is proposed for grading whether support tickets were resolved, meaning the customer's problem was fixed and not merely answered. The gold set is built from resolved and unresolved tickets in a simulated support environment. It includes near misses, such as replies that address the wrong account or promise a refund that was never issued, and outputs from an agent that had been tuned using an earlier version of the same judge. Two support experts label each ticket blind, with a third adjudicating. Items where the original customer request was itself ambiguous are set aside.

The judge's false-accept rate on near misses turns out to be far higher than on clear failures, and the on-policy stratum shows that the tuned agent has learned to write confident closing messages the judge rewards. The exploit audit confirms it: an attacker reaches a high acceptance rate on unresolved tickets with modest effort. The grader card grants the judge dashboard use only. Billing and training rewards for this class are routed to a deterministic check of the ticketing system's state, with the judge reserved for the semantic quality of replies. None of this is a result. It shows how the protocol turns abstract concerns into specific eligibility decisions.

Cost and staffing

The dominant cost is expert time. Gold sets can be amortized: the same expert-labeled items certify successive versions of a verifier, and certify competing verifiers for the same class. Anchor subsets can be small enough to re-run cheaply on every version change. Sharing gold sets across organizations is attractive but risks leakage into training data. Sealing, canaries and access logging reduce that risk without eliminating it. Where expert time is scarce, the protocol prioritizes task classes by the consequence of a false accept, and certifies those first.

Drift monitoring

A certified verifier must stay certified. A frozen anchor subset of the gold set is re-run on every new verifier version and periodically on the current version. If error rates on the anchor shift beyond a pre-registered tolerance, the verifier's eligibility for training rewards and billing is suspended until it is recertified. Model judges served by external providers need particular attention, because the model behind a stable name can change.

Threats to validity

  • Expert error. Experts are fallible. Adjudication and agreement statistics bound the problem; they do not remove it.
  • Gold-set staleness. Agents improve and find new failure modes. Gold sets must be refreshed with on-policy samples.
  • Leakage. If gold items leak into an agent's training data or prompts, measured verifier performance may stay stable while real performance changes. Gold sets are sealed and marked with canaries.
  • Class definition. Error rates vary within a task class. Classes that are too broad hide sub-classes where the verifier fails badly. Stratified reporting within classes is encouraged.

Relation to the benchmark program

Every track of Ethen VerifiedWork uses graders certified by this protocol and publishes their cards. The cost of verification and its effect on economic metrics are discussed in Cost Per Verified Outcome.

Limitations

This protocol has not been run on Ethen verifiers. Expert labeling is expensive, especially in specialized domains, and the gold-set sizes needed for tight bounds on rare errors may be impractical for some classes. Exploit audits find only what attackers think to try. Calibration methods for judges are evolving, and the protocol will need revision as they mature.

Conclusion

A verifier is an instrument, and instruments have error. This protocol measures that error where it matters: on hard cases, on the outputs optimized agents actually produce, under known biases and under attack. It then grants each verifier only the uses its measured reliability supports. Until a verifier has a grader card, its verdicts should inform people, not machines or invoices.

FAQ

How many gold items does a verifier need? It depends on the error rates that matter. 100 to 200 per task class is a starting point; safety-relevant classes need more, because rare errors require large samples to bound.

Why include outputs from optimized agents? Because a verifier's errors on outputs from agents tuned against it are what determine whether learning goes wrong. Random samples from untuned agents understate them.

Can one verifier be certified for everything? No. Certification is per task class and per use.

References

  1. Skalse, J. et al. (2022). Defining and Characterizing Reward Hacking. arXiv:2209.13085. https://arxiv.org/abs/2209.13085
  2. Pan, A. et al. (2022). The Effects of Reward Misspecification. arXiv:2201.03544. https://arxiv.org/abs/2201.03544
  3. Gao, L., Schulman, J., Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760. https://arxiv.org/abs/2210.10760
  4. Zheng, L. et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685. https://arxiv.org/abs/2306.05685
  5. Wang, P. et al. (2023). Large Language Models are not Fair Evaluators. arXiv:2305.17926. https://arxiv.org/abs/2305.17926
  6. Panickssery, A. et al. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076. https://arxiv.org/abs/2404.13076
  7. Thakur, A. S. et al. (2024). Judging the Judges. arXiv:2406.12624. https://arxiv.org/abs/2406.12624
  8. Zhu, Y. et al. (2025). Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv:2507.02825. https://arxiv.org/abs/2507.02825
  9. Angelopoulos, A. N. et al. (2023). Prediction-Powered Inference. arXiv:2301.09633. https://arxiv.org/abs/2301.09633
  10. Saad-Falcon, J. et al. (2023). ARES. arXiv:2311.09476. https://arxiv.org/abs/2311.09476
  11. Jung, J. et al. (2024). Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement. arXiv:2407.18370. https://arxiv.org/abs/2407.18370
  12. Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1). https://doi.org/10.1177/001316446002000104
  13. Hanley, J. A., Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? JAMA 249(13):1743–1745. https://doi.org/10.1001/jama.1983.03330370053031
  14. Lambert, N. et al. (2024). RewardBench. arXiv:2403.13787. https://arxiv.org/abs/2403.13787

Each publication states its evidence status. Designs, protocols, and proposals report no measured results.

  • Product

    What Ethen Is Doing to Make AI Outputs Easier to Verify

    The practical answer to how to verify AI output is to check claims, not paragraphs: split an answer into individual statements, open each source, find the exact sentence that supports each statement, make sure the support comes from independent sources, and keep anything you could not confirm labeled as unconfirmed. That works, but it is slow, which is why most AI output goes unchecked. Ethen's approach is to make verification cheaper by building it into the output. Research reports attach a status and the supporting passages to each claim, and say plainly what the check cannot do. Model facts without complete provenance show "Unknown" instead of a guess. Agent work is marked complete only after a separate verifier checks it. Verification capacity is reserved before work is spent. And releases come with scoped evidence. This article explains each mechanism, its limits, and what you should still check yourself.

  • Company

    What Ethen Research Lab Is Exploring Beyond AI Products

    Ethen Research Lab research areas reach beyond any single product. The Lab is organized into eight published programs — Evaluation and Verification; Trust and Accountable AI Work; Adaptive Intelligence; Context, Skills and Transfer; Model Intelligence and Faros; Data and Learning Systems; Enterprise and Sovereign AI; and AI Security — connected by one thread: how AI systems can learn from work that can be checked. Some questions feed directly into products. Others look further out: how to measure whether automated graders can be trusted, whether skills survive when the underlying model changes, whether organizations can improve AI without exporting their data, whether agents can predict the effects of their actions, and whether automated research can keep hypotheses separate from confirmed findings. Almost all of this work is published as proposals, protocols and benchmark designs; very little is results. This article describes the questions honestly, without claims the evidence cannot support.

Explore this topic

Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.