Skip to content

EthenEthenEthen

Checking the Evidence in an AI Research Report

A claim-by-claim method for reading source coverage, contradictions, and lexical matches without mistaking them for truth.

A claim-by-claim method for reading source coverage, contradictions, and lexical matches without mistaking them for truth.

An AI research report looks finished: sentences, citations, confident transitions. None of that shows whether each sentence is backed by its sources. AI research report verification starts one level down. Split the report into checkable claims, then ask what each source passage contributes. This guide shows a durable way to do that, using one deterministic local claim checker as a concrete reference. The checker counts word overlap between report sentences and source sentences. It does not judge meaning, and its own output says so.

What "verification" means here

Here, verification means claim-level evidence matching, not a verdict that a report is true. The reference implementation takes a report string and a list of sources, splits both into sentences, and records which source sentences share enough words with each report sentence. Each report sentence becomes a claim record with a status, a confidence label, matching passages, supporting and contradictory passage counts, and an explicit limitations note.

That note is fixed text: deterministic lexical passage matching is not a certified semantic verifier. The verifier record agrees: kind deterministic local, certified flag false, receipt a local lexical version marker. A high overlap count means shared vocabulary, not shared meaning. A supported status means at least two supporting passages from at least two independent sources matched lexically — not that the claim survived expert review. Lexical checks are repeatable and cheap to re-run. They catch missing citations, single-source claims, and passages pulling in opposite directions. They cannot catch a fluent sentence that restates a premise incorrectly or a source that is itself wrong.

Start with atomic claims, not paragraphs

One paragraph can hold three claims with three different evidence states. The checker therefore splits the report into sentences on sentence-ending punctuation followed by whitespace, trims each sentence, and keeps only sentences with at least two significant terms. Each survivor gets a sequential identifier such as claim-1, claim-2, and so on. Fragments below the two-term floor are dropped before verification begins.

Significant terms have a precise definition. Text is lowercased, scanned for alphanumeric runs of at least three characters, deduplicated, and filtered against a small stopword list: the, a, an, and, or, of, to, in, on, for, with, is, are, was, were, be, as, by, from, that, this, and it. Everything remaining counts as a wanted term for that claim.

Three habits follow from that definition. First, check one sentence at a time; if a sentence joins two facts with "and," verify each half separately in your notes even though the checker scores the sentence as one unit. Second, expect transitions and headings to vanish from the claim list — dropped text is unverified by construction, not confirmed. Third, watch sentences whose meaning lives in short tokens: two-letter abbreviations and single distinctive words contribute less than expected because only runs of three or more characters count, and repetition collapses to one term.

Source coverage: count independent support

For each claim, the checker loops over every supplied source, takes full text when present and otherwise the snippet, splits that text into sentences, and scores each source sentence against the claim's wanted terms. Overlap is the count of wanted terms also present among that sentence's terms. A source sentence becomes evidence only at or above the maximum of two terms and the ceiling of thirty-five percent of wanted terms. A five-term claim needs two matching terms; a ten-term claim needs four.

Each qualifying sentence is stored as a passage record: source identifier, source URL and title, a location marker of the form sentence-N, the excerpt truncated to 500 characters, a polarity of support or contradiction, and the overlap count. The claim then summarizes three counts: supporting passages, contradictory passages, and independent sources among the supporting passages. Independence means distinct source identifiers among supporting passages only. Two matching sentences from one document are two passages but one independent source.

The status rule is exact:

  • Contested: at least one supporting and at least one contradictory passage.
  • Supported: at least two supporting passages from at least two independent sources, with no contradictory passage.
  • Weakly supported: at least one supporting passage but below the supported bar, with no contradictory passage.
  • Unsupported: no supporting or contradictory passages.

A fifth status value, not assessed, exists in the type but the matching function never assigns it.

Use this as a coverage checklist per claim. How many independent sources support it? One source is weak by definition, however long the excerpt. Is support spread across genuinely different sources? Does any passage contradict it — a single contradictory match forces contested even beside several supporting passages. Is the best match full text or a snippet fallback? An empty source contributes nothing.

Note what the rule never requires: sources agreeing with each other beyond shared vocabulary, sources being primary or authoritative, numbers and causal connectors matching. Those judgments are yours. The counts tell you where to look: unsupported claims need sources, weakly supported claims need a second independent source, contested claims need a human to read both sides.

Contradictory passages: one flag changes the verdict

Polarity comes from one case-insensitive word test on each qualifying source sentence. A sentence is a contradiction when it contains any whole word from this list: not, no, never, false, incorrect, decline, declined, declining, contradicts, contradicted, contradiction. Every other qualifying sentence is support. There is no middle polarity and no per-sentence confidence.

Two consequences matter. First, contested is easy to trigger and always worth reading: one "not" or "no" in an overlapping sentence flags the claim. That is triage behavior — surface tension rather than resolve it. Second, the flag is lexical, not logical. "Not only X but also Y" contains a negation word while agreeing; "results differed across runs" rebuts without any listed word. Both will be mislabeled by polarity alone.

Read each contradictory passage in its source document, because the excerpt is capped at 500 characters and the location is only a sentence number. Ask whether the negation attaches to the claim's core predicate: "the method did not fail" supports a success claim despite the negation, while "the method did not replicate" opposes it. Check whether support and contradiction come from the same source — one hedging document — or across sources, which can signal genuine disagreement. Record your resolution outside the checker: the claim's review history starts with one machine judgment entry stamped with the current time, and the event type allows a manual override, but the matching function only ever emits the machine judgment. Your override is a separate human step.

Never read the contradictory count as a disagreement meter. It counts lexically matching sentences containing a negation-family word. Two hits may quote one sentence twice, and zero hits may hide a polite rebuttal phrased without any listed word.

The limits of lexical checks

Every number in the claim record traces to shared terms, so every limit of term matching limits the verdict. Four recur often enough to memorize.

Threshold effects come first. Short claims are brittle and long claims demanding. A two-term claim is supported by any source sentence containing both terms — thin when the terms are common. A twelve-term claim needs five matching terms, which a fair paraphrase may miss. When a claim you believe is well sourced returns unsupported, count its wanted terms by hand: lowercase, keep runs of three or more alphanumeric characters, drop the stopwords, deduplicate. If the threshold is the problem, narrow the claim into smaller checkable sentences rather than silently lowering the bar.

Vocabulary mismatch is second. There are no synonyms, no stemming beyond literal runs, and no entity resolution. "GPU rental," "accelerator hire," and a vendor product name are different terms unless they share literal runs, so a renaming claim scores low even when a human would call it supported. The reverse also holds: reusing the source's nouns while reversing its verbs can score high overlap while saying the opposite. The stored overlap maximum — the largest single-sentence overlap in the claim's evidence — measures the best vocabulary hit, not the closest meaning. A high maximum beside a contradiction flag is the classic pattern: same words, opposite direction.

Stopwords and short tokens are third. Negation words illustrate the trap: "not" and "no" drive polarity yet never count toward overlap as terms. Tense, modality, and hedging ("may," "can," "often") fall below the length floor or add threshold burden without discrimination. Read modal claims by hand even when they score supported.

Truncation and splitting are fourth. Passages keep 500 characters, long sentences lose their tails, and sentence splitting can separate a caveat from its claim. The excerpt is a pointer, not a citable quotation — open the source.

Beneath all four sits the limit the code states outright: lexical overlap is not a semantic truth guarantee. Shared words do not establish that the source is correct or that the report's inference follows. Two independent sources can repeat one incorrect announcement; the checker counts two identifiers and marks the claim supported, because independence means distinct source identifiers, not independent investigation or data. That is why lexical results must never be presented as research quality measurements. They are coverage measurements: how many supplied passages share vocabulary with each sentence, and whether any contain negation-family words.

Freshness and retrieval gaps

Each source may carry a retrieved-at timestamp, and the claim record counts sources lacking one or carrying an unparsable date. That stale-source count is the only freshness signal. It does not compare dates against the claim, rank newer sources higher, or discount old passages. A years-old retrieval and today's retrieval contribute equally if their identifiers differ.

Treat the count as a prompt. If every source is undated, you have no retrieval timeline — ask when each was fetched and whether the page may have changed. If the claim is time-sensitive, such as pricing, availability, or version behavior, an undated retrieval is a reason to re-fetch, not to trust the excerpt. If a source supplied only a snippet because full text was unavailable, remember the checker scored the snippet: coverage against someone else's summary, not the document.

One input boundary deserves emphasis. The code comments state that retrieved text is treated as data to be scored, never as instructions to follow; only delimited sentence content enters the overlap computation. That scopes the checker, not any larger research pipeline. Confirm separately how retrieved content flows into prompts, tools, and actions elsewhere.

What the job layer proves, and what it does not

Claims are checked after research runs, and research runs inside durable jobs. The companion source here is the canonical research worker handler — worth understanding because readers confuse reliable execution with reliable findings. The handler delivers the former and says nothing about the latter.

Three handler objects share one durable path: a research handler, a research-dispatch handler delegating to it, and a deep-research handler with its own job kind. All three build a platform job repository and durable job service, send a lease heartbeat before dispatch, and route through a shared durable research worker. There is deliberately no second job platform: one durable jobs store plus the shared service. Before any external provider effect — the code names Exa at that boundary — each handler records a provider-dispatched marker keyed by job identifier plus mode or run identifier, so a retried job suppresses duplicate external calls. Failures distinguish lease loss from other errors, and only lease loss is retryable. Completion is recorded by the worker host as the canonical completed event; custom handler-level event names are rejected by the durable job event log, and the comments warn that invented names previously masked real terminal states behind persistence errors.

All of that concerns execution durability: leases, heartbeats, idempotent dispatch, retry policy, canonical completion. None of it speaks to source relevance, inference validity, or report quality. A job can complete cleanly with a full event history while every claim is weakly supported or unsupported. A lease-lost retry says the worker lost its claim on the job, not that the research was wrong. Keep the layers separate: the job layer answers whether work ran once and terminated in a recorded state; the claim layer answers how many supplied passages share vocabulary per sentence. Research quality — sound questions, apt methods, honest interpretation — is a third layer neither path measures.

One architecture note frames any product claims. In the current target map, Chat is limited and Studio, Research, Designer, and Founder are separate target applications, with Founder as its own product, Computer as a Platform-only direction, Desktop owning local models, and Code spanning lightweight Chat, cloud Platform, and local Desktop. Target separation describes intended ownership, not proof that any migration or launch shipped. The approved marketing copy describes the Research direction as a product for investigating questions and organizing research work, separate from Labs — a direction statement, not availability evidence. Read this guide as a durable skill, not as a claim about any hosted product today.

A durable review checklist

Apply these steps to any AI research report, by hand or with tooling:

  1. Split the report into numbered sentences. Mark which carry factual load; only those can have evidence coverage.
  2. For each factual sentence, name its significant terms — nouns, verbs, distinctive phrases a supporting passage must contain. Fewer than two means too vague to verify as written.
  3. Find supporting passages and record source identifiers. Two passages from one source are still one independent source; seek two independent sources before calling a claim covered.
  4. Read every contradictory passage in full context. Decide whether the negation hits the claim's predicate and whether tension sits inside one source or across sources.
  5. Test paraphrase robustness. Would a fair rewording still share enough distinctive words to match? Support resting on one shared noun is thin.
  6. Check freshness and provenance. When was each source retrieved, full text or snippet, and are the two sources independent investigations or two copies of one announcement?
  7. Separate execution from evidence. A clean job log proves the pipeline ran; score evidence from passages, not run metadata.
  8. Record manual judgments. Note confirmed versus dismissed contradictions, and narrowed, re-sourced, or removed unsupported claims.

Worked example, clearly illustrative

This miniature report and its sources are invented only to show how statuses behave — not measurements of any system.

Report: (1) "The north depot rollout finished in March and cut night-shift overtime." (2) "Operators prefer the new scheduling board." (3) "The rollout did not increase daytime staffing costs."

Sources: A says the north depot rollout finished in March and night-shift overtime fell afterward. B says the rollout finished in March and overtime declined. C says operators never adopted the new board and daytime staffing costs were not reviewed.

Claim 1 shares rollout, finished, March, night-shift, and overtime terms with passages in A and B, with no negation word in those sentences: supporting passages from two independent sources, hence supported. Claim 2 shares operators, scheduling, and board terms only with C's "never adopted" sentence — the sole match is a contradiction, and with zero supporting passages the claim is unsupported, not contested. That asymmetry surprises readers: contested requires at least one passage on each side. Claim 3 shares staffing and costs terms with C's "not reviewed" sentence; pair that contradiction with a supporting daytime-costs sentence elsewhere and it becomes contested, requiring a human to decide whether costs were measured at all.

Now vary the inputs. If A and B quote one depot newsletter, independence still reads two — distinct identifiers, not distinct reporting. If B instead says "the northern hub deployment concluded in spring with fewer extra hours," a human sees paraphrase support but the checker sees almost no shared literal terms. If A's date is wrong, supported does not catch it. The table measures vocabulary coverage across supplied identifiers; your review supplies the rest.

When evidence is thin

Thin evidence is the normal early-draft state, and the labels give calm responses. Unsupported means the sentence is too vague, unaddressed, or phrased in different words: narrow it, add a source that discusses it, or split it until each half matches or clearly fails. Weakly supported means a start without coverage: find a second independent source or soften the sentence to what one source carries ("one field report describes" rather than a general claim). Contested means stop and read — quote both passages, check dates and scopes, and either scope the sentence to where it holds or state the disagreement.

Avoid two shortcuts. Do not reword the report to match source vocabulary purely to raise overlap; the score improves while the evidence stays flat and the sentence may drift from what you can defend. Do not stack sources sharing one origin; one more independent investigation beats five more copies. When no second source or scope fix exists, the honest status is unsupported and the honest edit is to cut or hedge.

Keep the limitations in view

End each review by restating what the table never did: assess meaning, weigh source quality, verify dates against reality, or measure research quality. Confidence tracks the status rule — supported maps to high, weakly supported to low, everything else including contested to unsupported — not your credence. The contradiction test is a fixed word list, independence is identifier inequality, freshness is timestamp presence. The receipt string and uncertified flag say exactly that.

Read each claim record as a work ticket: the sentence, the passages sharing its words, whether any contain negation terms, and what a human must still decide. That turns "this report feels thin" into "claims 3, 6, and 8 lack supporting passages; claims 2 and 5 are contested; claim 7 rests on one source" — fixable items with review history attached. Certification would need independent fact-checking against primary sources and domain review of methods, work the lexical checker never attempts. Keep those labors distinct and the evidence table earns its place: a fast, repeatable first pass telling the reviewer exactly where to look next.