Skip to content

EthenEthenEthen

What Ethen Is Doing to Make AI Outputs Easier to Verify

The practical answer to how to verify AI output is to check claims, not paragraphs: split an answer into individual statements, open each source, find the exact sentence that supports each statement, make sure the support comes from independent sources, and keep anything you could not confirm labeled as unconfirmed. That works, but it is slow, which is why most AI output goes unchecked. Ethen's approach is to make verification cheaper by building it into the output. Research reports attach a status and the supporting passages to each claim, and say plainly what the check cannot do. Model facts without complete provenance show "Unknown" instead of a guess. Agent work is marked complete only after a separate verifier checks it. Verification capacity is reserved before work is spent. And releases come with scoped evidence. This article explains each mechanism, its limits, and what you should still check yourself.

The practical answer to how to verify AI output is to check claims, not paragraphs: split an answer into individual statements, open each source, find the exact sentence that supports each statement, make sure the support comes from independent sources, and keep anything you could not confirm labeled as unconfirmed. That works, but it is slow, which is why most AI output goes unchecked. Ethen's approach is to make verification cheaper by building it into the output. Research reports attach a status and the supporting passages to each claim, and say plainly what the check cannot do. Model facts without complete provenance show "Unknown" instead of a guess. Agent work is marked complete only after a separate verifier checks it. Verification capacity is reserved before work is spent. And releases come with scoped evidence. This article explains each mechanism, its limits, and what you should still check yourself.

Key takeaways

  • Verify claims, not answers. An answer is many claims, each with its own evidence.
  • A citation link is not proof. Research has found many generated sentences are not fully supported by the sources cited.
  • Statuses direct attention. A claim marked contested or unsupported tells you where to look first.
  • Unknown is an honest value. A blank where evidence is missing beats a confident guess.
  • Completion needs a separate check. An agent saying "done" is a claim, not evidence.
  • Automation supports judgment; it does not replace it. Every mechanism here states its limits.

Why AI output is hard to verify

AI output is hard to verify because it is fluent. A well-written paragraph with citations looks finished, and nothing in its surface tells you which sentences are supported and which are not.

Research has measured the gap. A 2023 study of generative search engines, by Nelson Liu, Tianyi Zhang and Percy Liang, used human evaluators to check whether generated sentences were supported by the citations attached to them. On average, only about half of generated sentences were fully supported by their citations, and about a quarter of citations did not support the sentence they were attached to. Systems and models have changed since, but the lesson holds: the presence of a citation tells you little about whether the claim beside it is supported.

A broad survey of hallucination in natural language generation describes the same underlying problem from the model side: language models can produce content that is fluent but unfaithful to their sources or to the world. Better models reduce the rate. They do not remove the need to check.

The cost of checking is the real obstacle. If verifying an answer takes as long as producing it yourself, people stop verifying. So the design question is not "how do we tell users to verify?" but "how do we make verification fast enough that people actually do it?"

Five places Ethen builds verification in

Ethen applies the same idea — attach the evidence and say what it does not cover — in several products. Figure 1 shows five published mechanisms.

Five bands of published verification mechanisms: claim statuses in research reports (highlighted), unknowns in model facts, verifier-gated completion of agent work, reserved verification budgets, and scoped release certificates.
Figure 1. Different products, one idea: attach the evidence and say what it does not cover.

1. Research reports: a status for every claim

When Ethen produces a research report, the verification step works at the level of individual claims. The report is split into sentences, and each one is matched against passages from the sources that were gathered. Each claim receives a status, the passages that match it, and counts of supporting and contradicting passages. We explain the method in detail in Checking the Evidence in an AI Research Report. Figure 2 summarizes the published rules.

Table of four claim statuses with their rules and suggested reader actions: supported (highlighted), weakly supported, contested and unsupported; legend notes matching is lexical, not semantic.
Figure 2. A status tells you where to spend your checking time, not whether a claim is true.

A claim is supported when at least two passages from at least two independent sources match it and none contradicts it. It is weakly supported when there is some support but less than that. It is contested when at least one passage supports it and at least one contradicts it. And it is unsupported when nothing matches.

The limits are stated as clearly as the rules. The matching is lexical: it counts shared words between report sentences and source sentences. It does not understand meaning, and the check describes itself as not a certified semantic verifier. A sentence that reuses a source's nouns while reversing its meaning can score well; a correct paraphrase in different words can score poorly. Two sources repeating the same wrong announcement count as two sources. That is why the statuses are best read as triage — where to look first — rather than verdicts.

That triage is still valuable. A reader with five minutes can skip the claims marked supported, read the contested ones carefully, and treat the unsupported ones as unverified. That turns an all-or-nothing chore into a targeted review.

2. Model facts: Unknown instead of a guess

When Ethen shows information about AI models — benchmark results, capabilities, specifications — a value appears only when its provenance is complete. If the source, method, version or freshness is missing, the cell shows "Unknown". We explain the reasoning in Showing Unknowns in Ethen Model Intelligence, and the broader principle in Why Ethen Shows What It Knows—and What It Doesn't.

For verification, an honest unknown is a gift: it tells you exactly which facts you would need to find elsewhere, instead of hiding a guess among real values.

3. Agent work: completion checked by someone other than the worker

When an Ethen agent carries out a multi-step task, the published design for missions routes completion through a verification step, and the verifier stands outside the executor. The part of the system that does the work proposes that it is done; a separate check decides. When the outcome of an action cannot be determined — a request timed out, a confirmation never arrived — the state is recorded as unknown rather than quietly treated as success or failure. We describe this in Making Mission Completion Depend on Evidence and the general principle in What "Done" Should Mean for an AI Agent.

4. Budgets: reserve capacity for checking first

Verification costs something: model calls, compute, time. A system that spends its whole budget doing the work has nothing left to check it. Ethen's published mission design reserves verification capacity before the executor spends, so that checking cannot be squeezed out by a task that ran long. The mechanism is described in Reserving a Budget for Verification.

5. Releases: scoped evidence, not blanket claims

Ethen's release certificates record dated, scoped evidence about a specific release — which checks passed, which boundaries were confirmed and what was not tested. They are the product-operations version of the same idea. See What a Release Certificate Actually Proves at Ethen.

Why verification is a product problem, not only a user problem

The usual advice about AI output puts the burden on users: double-check everything. That advice is correct and mostly ignored, because it asks people to spend the time the AI was supposed to save. We think the responsibility should be shared, and that most of it belongs with the product.

A product can make verification cheaper in ways a user cannot. It can keep track of which source passage each claim came from, instead of attaching citations afterwards. It can refuse to fill in a value it does not know. It can route work through an independent check before calling it done. It can reserve the budget for that check up front. And it can present all of this in a way that tells the user where to look, rather than leaving them to audit everything equally.

There is a second reason. As AI systems take on longer tasks — research spanning dozens of sources, agents acting across several tools — the amount of output grows faster than anyone's capacity to review it. Without verification built into the product, checking simply stops happening. Building it in is the only way review keeps pace with output.

This is also why verification connects to so much else at Ethen: approvals before consequential actions, honest handling of unknown outcomes, and evidence records for completed work. Each makes a different kind of claim checkable.

How to verify any AI answer yourself

Ethen's mechanisms make verification cheaper, but the same habits work with any AI tool, including ours. Figure 3 summarizes them.

Five-step checklist for verifying an AI answer: split into claims, open the source, find the exact supporting sentence (highlighted), check source independence, and mark what remains unverified.
Figure 3. Works for any AI tool, including Ethen.

Split into claims. An answer is a bundle of statements. Pick the ones that matter for your decision and check them individually.

Open the source. A link that resolves is not evidence. The page may not say what the answer claims.

Find the sentence. Locate the specific words that support the claim. If you cannot find them, the claim is unsupported by that source, however authoritative the source.

Check independence. Several articles repeating one press release are one source. Look for support from independent origins — different data, different investigators.

Mark what is left. Some claims will remain unconfirmed. Keep them labeled that way in anything you pass on.

What automated verification can and cannot do

It is worth being precise about the role of automation, because overclaiming here would undermine the whole point.

Automated checks are good at coverage. They can tell you which claims have matching passages, which rely on a single source, and where sources pull in opposite directions. They are repeatable and cheap to rerun.

Automated checks are weak at meaning. Lexical matching cannot tell whether a paraphrase is faithful, whether a negation attaches to the right part of a sentence, or whether a source is itself wrong. More sophisticated checkers that use language models to judge support bring their own errors.

Verifiers need verifying. If a language model is used to judge whether another model's output is correct, the judge's reliability has to be measured too. Ethen Research Lab has published a protocol for exactly that question — How Should We Measure the Reliability of LLM Verifiers? — which has not yet been run. A related methods paper, Evaluating the Evaluators: Reward Integrity for AI Agents, discusses how checks can be gamed. Neither reports results about Ethen's verifiers.

The honest position is that automated verification shifts human effort from checking everything to checking what matters. It does not remove human judgment, and Ethen does not present it as doing so.

A worked example

The following example is illustrative. A product manager asks Ethen for a short research report on whether a new data-privacy regulation applies to their company's analytics tool.

The report comes back with eight claims. Five are marked supported, each with passages from two or more independent sources: the regulation's official text and two law-firm summaries. One is marked contested: one summary says a particular exemption applies to small companies, while the official text appears to limit it. One is weakly supported, relying on a single blog post. One is unsupported: a statement about an enforcement date that no gathered source mentions.

The manager skims the five supported claims, reads the contested passages in full and finds that the law-firm summary was out of date, looks for a second source for the weakly supported claim, and removes the unsupported enforcement date. Total checking time: about fifteen minutes, instead of re-researching the question from scratch. The remaining judgment — whether the regulation applies to their specific product — goes to their legal adviser, with a report that shows exactly what was and was not confirmed.

What we are working on next

Several directions are open. Better handling of paraphrase, so that faithful restatements are recognized as supported. Clearer presentation of claim statuses in the interface, so the triage is visible at a glance. Extending claim-level evidence from research reports to other kinds of output, such as summaries and drafts that rely on documents you provided. And measuring verifier reliability, so that statements about how well checks work rest on data. These are directions, not announcements; we will describe them when they ship.

Tradeoffs and limitations

Verification adds cost and time. Matching claims to passages and running separate checks makes outputs slower and more expensive to produce. We think that is the right trade for work people rely on.

Statuses can create false confidence. A "supported" label based on word overlap is not a guarantee. We state that limit wherever the status appears, and readers should still spot-check.

Some claims cannot be verified by sources. Judgments, predictions and recommendations need human evaluation, not passage matching.

Mechanisms are at different stages. The research-report checks, model unknowns, evidence-gated mission completion and verification budgets are published as mechanisms with stated test limits. They are not claims of measured accuracy.

FAQ

How can I check whether an AI answer is correct? Split it into claims, open the cited sources, find the exact sentence supporting each claim, check that support comes from independent sources, and keep unconfirmed claims labeled.

Are AI citations reliable? Not on their own. A 2023 study found only about half of generated sentences were fully supported by their citations. Always open the source.

How does Ethen help verify AI outputs? By attaching evidence and a status to claims in research reports, showing unknown values honestly, checking agent work with a separate verifier before marking it done, reserving verification capacity, and publishing scoped release evidence.

Does Ethen guarantee its outputs are correct? No. Its checks make verification faster and more targeted, and each states its limits. Human judgment remains necessary for important decisions.

What does "contested" mean in an Ethen research report? At least one source passage supports the claim and at least one contradicts it. Read both passages in full before relying on the claim.

References

  1. Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. Findings of EMNLP 2023. arXiv:2304.09848. https://arxiv.org/abs/2304.09848
  2. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12), 1–38. https://doi.org/10.1145/3571730
  3. Ethen Research Lab (2026). How Should We Measure the Reliability of LLM Verifiers? Research protocol; not yet run. https://upcube.ai/resources/research/measuring-llm-verifier-reliability
  4. Ethen Research Lab (2026). Evaluating the Evaluators: Reward Integrity for AI Agents. Methods paper. https://upcube.ai/resources/research/reward-integrity
  5. Ethen Blog. Checking the Evidence in an AI Research Report. https://upcube.ai/blog/checking-the-evidence-in-an-ai-research-report