Skip to content

EthenEthenEthen

What Makes an AI Agent Job Verifiable

An agent job is verifiable when someone other than the agent can confirm what was supposed to happen, what actually happened, and what remains unknown. This guide turns that idea into a checklist you can apply to any system.

An agent job is verifiable when someone other than the agent can confirm what was supposed to happen, what actually happened, and what remains unknown. This guide turns that idea into a checklist you can apply to any system.

Most demos of AI agents end at the moment the agent says it is done. That is the least informative moment in the whole run. A chat answer can be skimmed and judged on the spot, but an agent job unfolds over time: it plans, acts on external systems, gets interrupted, retries, and sometimes leaves effects nobody has confirmed. If the only record of success is the agent's own summary, you have a story, not a result.

Verifiability is the property that closes this gap. A verifiable job defines its outcome before it starts, exposes points where a person can intervene, recovers from interruptions without inventing facts, and produces evidence that an independent check can evaluate. None of this requires trusting the agent's self-assessment. Each element below is something you can ask a vendor about, look for in documentation, or require in your own build.

The checklist is organized in four parts: outcomes, intervention, recovery, and evidence. It draws on concrete mechanisms found in inspected local sources — a mission state machine with guarded transitions, a durable reconciler that refuses to guess at unknown effects, and an independent audit of that system's remediation stage — while staying general enough to apply to any agent platform. Illustrative examples in this article are hypothetical; fixtures and walkthroughs cannot establish autonomous completion rates.

1. Outcomes: define done before the job starts

A job that cannot describe its own completion cannot be verified. The first set of questions pins down what "done" means and who is allowed to declare it.

Does the system model the job's lifecycle as explicit states? Look for a named set of states with legal transitions between them, not a free-text status field. In the inspected mission state machine, a mission moves through sixteen uppercase states (from CREATED through CONTRACTING, PLANNING, READY, RUNNING, VERIFYING, and on to a terminal state) while each task moves through nine (PENDING to SUCCEEDED, FAILED, or CANCELED). Every edge is enumerated in a transition table, and anything not listed is rejected. This matters because explicit states make two guarantees possible: illegal jumps are impossible by construction, and every job's history reads as a legible sequence a reviewer can follow.

Are terminal states truly terminal? A COMPLETED job that can quietly reopen is a reporting hazard: downstream steps, billing, and audits all assumed finality. Ask whether completion, failure, and cancellation are one-way doors. The inspected reducers enforce this directly — once a mission reaches COMPLETED, FAILED, or ABORTED, no further transition is accepted, and the same holds for terminal task states. Stale writes against old revisions are rejected rather than applied. When you evaluate a system, try the adversarial question: what happens if a delayed retry or a second worker tries to move a finished job? "It fails closed" is the answer you want.

Is executor completion separated from certified success? An agent finishing its own steps is not the same as the job succeeding. The inspected design makes this distinction structural: a task whose executor finished moves to VERIFYING, and the step into SUCCEEDED additionally requires a reference to an independent check. Executor completion alone never unlocks dependents. Ask any platform: can a task be marked successful by the same component that executed it, or does success require a separate check with its own reference? If the executor can self-certify, dependents and dashboards inherit its blind spots.

Is the vocabulary honest about approximations? Real systems project rich internal states onto simpler external labels — "queued," "running," "paused" — for dashboards and integrations. That projection should be documented as an approximation for interop, never presented as proof that the internal job is the simple label. Check that unmapped states surface as errors rather than silently collapsing into the nearest familiar word.

Outcome checklist, summarized:

  • Named lifecycle states with an explicit transition table.
  • Terminal states that can never reopen; stale revisions rejected.
  • A verification step between "executor finished" and "succeeded."
  • Documented, honest projection of internal states onto display labels.

2. Intervention: a person can step in at any point

Long-running jobs encounter situations no plan covers: ambiguous approvals, blocked dependencies, changed instructions, or simply a human who wants to stop the run. Verifiability requires that intervention is always possible and always leaves a legible trace.

Can approval, pause, block, and cancellation interrupt any active state? The strongest pattern is an "any-nonterminal" family of edges: from every unfinished state, the job can move to awaiting approval, blocked, paused, canceling, or failed. In the inspected state machine, exactly this family exists — approval requests, blocks, pauses, cancellation, and fatal failure can interrupt any active mission state, with narrowly documented exceptions (never a self-loop, never new side-channels from inside cancellation, never out of a terminal state). Cancellation itself carries a closed set of exits so that a canceling job quiesces into a resolved outcome instead of hanging. When you evaluate a product, ask for its interruption matrix: from a running step, which human actions are legal, and where does each one land?

Are approvals bound to specific actions? A blanket "approved" flag is weak evidence. Stronger systems bind each approval to the specific action it authorizes — the action's identity, the run it belongs to, and the conditions under which the approval expires or is consumed. The independent audit describes missions-side human-approval binding for consequential dispatch, alongside a policy layer that classifies actions as allowed, approval-required, or blocked. Ask: if the plan changes after I approve step three, is my approval still valid for the new step three? It should not be.

Does intervention preserve the record? Pausing, blocking, or taking over a job must not destroy the trail. State transitions driven by intervention should be recorded with the same durability as autonomous ones, including who intervened and what was pending. A takeover that erases the pre-takeover history trades one problem for another: the job may finish, but nobody can reconstruct what the agent did before the human arrived.

Is authority fenced per action? Intervention is only meaningful if the agent cannot silently expand its own permissions between checkpoints. The inspected invariant is that authority shrinks and never silently widens: unknown or malformed authority is denied, cached permissions are keyed to a digest that is recompiled when the contract changes, and amending the contract invalidates outstanding approvals and proposals. Ask whether each dispatch re-checks the current grant, kill-switch, and budget — or whether a grant checked at planning time is trusted for the whole run.

Intervention checklist, summarized:

  • Approval, pause, block, cancel, and fail reachable from any unfinished state.
  • Approvals bound to specific actions, with expiry and single-use consumption.
  • Intervention recorded durably, including takeover and handback.
  • Authority re-checked at every dispatch; contract changes invalidate old approvals.

3. Recovery: unknown outcomes stay unknown until observed

This is the part most agent demos skip entirely. Real jobs get interrupted: workers crash, networks drop, leases expire, external systems respond ambiguously. After an interruption, some action effects are genuinely unknown — the request may or may not have gone through. How a system handles that uncertainty is the clearest single test of its verifiability.

Does the system distinguish unknown from failed? A failed action is known: it did not take effect, and retry or compensation can proceed. An unknown action is unresolved: retrying blindly may double-apply it, and declaring failure may be wrong. The inspected architecture treats these as different fenced states. Unknown effects are preserved across restarts, no timeout ever flips an unknown effect to a decided one, and final completion is blocked while any unknown effect remains. Ask the blunt question: after a crash mid-action, does your system retry, declare failure, or retain an explicit unknown until it re-observes the world? Only the last answer is verifiable.

Is there a durable reconciler with leased claims? Recovery should be a deliberate loop, not hopeful re-execution. The inspected reconciler works as a serve loop over recorded reconciliation obligations: it claims each expired or open obligation with a scoped lease, observes the effect exactly once by a stable key, and resolves it authoritatively. It is stateless across restarts — all truth comes from committed rows, never from process memory — and one bad row never kills a sweep: per-obligation errors are counted and skipped. Losing a lease race to another worker is skipped silently rather than counted as an error. When evaluating a platform, ask: where is the list of unresolved obligations stored, how are concurrent recovery workers prevented from double-resolving, and what survives a full restart?

Does recovery re-observe rather than assume? After time passes, after failures, after human edits, and after external changes, prior observations go stale. The inspected invariant requires re-observation across all of these: stale observations, stale world state, stale revisions, and stale generations are all rejected at the relevant check, and resume paths explicitly flag when fresh observation is needed before dispatch continues. A system that trusts a pre-crash snapshot after a human edited the world during the outage is building on sand.

Does the reconciler refuse to invent evidence? This is the honesty test. The inspected reconciler's lookup providers are injected, and at the audited stage the only live implementation was a test fake — there was no production provider adapter yet. In that situation the wired serve loop observes "unknown" and retains every obligation rather than manufacturing receipts. Lookup failures and conflicts likewise resolve to unknown, never to a guessed success or failure. Ask any vendor what their recovery path does when the external system cannot confirm an effect. Retaining an explicit unknown is correct; synthesizing a plausible receipt is fabrication, however convenient.

Is replanning bounded? Recovery often involves replanning, and unbounded replanning is a liveness risk — a job that re-plans forever never finishes and never cleanly fails. Look for explicit bounds (a maximum number of replans, generation fencing so old plans cannot overwrite new ones) and tests showing that carried-over progress survives a replan. The audited system bounds replanning by count plus generation fencing, with a noted suggestion to add hysteresis if replanning ever oscillates — an honest statement of a limit rather than a claim of perfection.

Recovery checklist, summarized:

  • Unknown and failed are distinct states; timeouts never decide unknowns.
  • A durable, leased, restart-safe reconciliation loop over committed obligations.
  • Mandatory re-observation after time, failure, human edits, and external change.
  • Unknown retained when effects cannot be confirmed — never invented.
  • Bounded replanning with generation fencing.

4. Evidence: completion requires independent proof

Outcomes, intervention, and recovery all produce the raw material. Evidence is what turns that material into a verdict someone can trust — ideally someone with no stake in the agent's success.

Is verification structurally independent? Independence is not a promise; it is an architecture. The audited system runs verification as a separate service with its own process, image, role, and queue: the verifier holds read privileges the executor lacks, re-hashes evidence bytes itself, and only the verifier role can record results and finalize missions. Executor attempts to finalize are denied and tested. Deterministic oracles back the checks, and appeal paths distinguish human acceptance from verifier verdicts. Ask: which component can mark this job complete, what credentials does it hold that the executor does not, and is that separation tested or merely documented?

Does completion check coverage, freshness, and contradiction? A verifier that only checks "did the executor attach something" adds little. Stronger verification confirms that the evidence covers the job's requirements, that observations are fresh enough to trust, and that contradictory evidence blocks completion rather than being averaged away. The audited checks include coverage analysis with explicit gaps, freshness windows, and a contradiction register where high-severity contradictions fail the job. Ask what happens when two pieces of evidence disagree, and when the newest observation is older than the freshness window. "It fails with a named reason" beats "it proceeds with a warning" for anything consequential.

Are evidence references non-empty and typed? Small gates catch large classes of error. In the inspected state machine, submitting a task for verification requires a non-empty set of evidence references, and entering SUCCEEDED requires a non-empty independent-check reference — empty strings and missing fields are rejected with a named error. These look like trivial validations, but they are what prevent "verified with no evidence" from ever being representable. Ask whether your system's data model can even express evidence-free success. If it can, the gate is policy, not structure.

Is the evidence trail reconstructable? After the fact, a reviewer should be able to rebuild the job's story from committed records: states visited, approvals consumed, observations made, unknowns retained and later resolved, and the verifier's verdict with its inputs. Durable event stores and tamper-refusing restores support this; summaries alone do not. Note the deliberate choice in the audited system: compaction is forbidden by design, because summaries must never become authoritative over the records they summarize. Ask whether the audit trail is the primary record or a derived view — and what happens to history when storage is compacted.

Evidence checklist, summarized:

  • A separate verifier with distinct privileges; executor cannot finalize.
  • Coverage, freshness, and contradiction checks with named failure reasons.
  • Non-empty evidence and check references enforced by the data model.
  • Reconstructable history from committed records; summaries never authoritative.

Putting it together: an illustrative walkthrough

Consider a hypothetical job: "reconcile last month's travel expenses and file the report." It is illustrative only — a walkthrough shows how the checklist applies, not how often real jobs succeed.

Before the run, the job contract states the outcome: every receipt matched to a ledger line, exceptions listed with reasons, and the report filed under the company's current policy version. The lifecycle states are explicit, terminal states are one-way, and the executor finishing its matching pass moves the job to verification rather than success. That covers outcomes.

Mid-run, the agent flags three receipts above the auto-approve threshold. The job moves to awaiting approval — reachable from its active state — and each approval is bound to a specific receipt and amount. While waiting, the finance lead pauses the job to update the travel policy; the pause is recorded, and the policy change invalidates the earlier approvals, which must be re-requested against the new rules. That covers intervention.

During filing, the worker crashes after submitting the report but before recording the confirmation. On restart, the reconciler claims the open obligation under a lease and looks up the filing by its stable idempotency key. The finance system is unreachable, so the effect stays explicitly unknown — the job cannot complete, and no receipt is synthesized. When connectivity returns, the lookup confirms the filing went through exactly once, and the obligation resolves. That covers recovery.

Finally, an independent verifier — holding read access the executor never had — re-checks the evidence: every receipt covered, observations fresh, no contradictions between the ledger and the filed report, approvals valid under the current policy. Only then does the job reach its terminal success state, with a history any auditor can replay. That covers evidence.

Notice what the walkthrough never claims: it says nothing about how often such jobs succeed unattended, how accurate the receipt matching is, or whether any particular product offers this flow today. Those are empirical and product-availability questions, and this checklist cannot answer them. It can only tell you what to look for.

What verifiability does not prove

A checklist this structural has sharp limits, and stating them is part of using it honestly.

First, fixtures and deterministic walkthroughs cannot establish autonomous completion rates. The audited evidence reported a deterministic bench run of fifteen cases and thirty-three trials, all passing — honestly labeled as deterministic pure checkers with no providers, network, or generated code involved. Live-model trials and deferred bench families were explicitly out of scope. That is the right way to report such results, and it means exactly what it says: the mechanism behaved correctly in its tested scope, not that agents complete real jobs at any particular rate. Whenever someone cites a pass count at you, ask what the checkers actually exercised.

Second, a dated audit proves only its scope at its date. The independent exit audit referenced here is a September 2026 Stage-0 assessment against a specific contract: it passed Stage-0 exit while deferring live sandbox use, browser takeover trials, per-connector production reconcilers, model-judge calibration, and continuous autonomy to later gates. Those deferrals are classified follow-ups, not hidden failures — but they mean the audit cannot support claims about capabilities it explicitly deferred. Treat every audit, including this reference, as a snapshot with boundaries.

Third, mechanism is not availability. The reconciler's honest-unknown behavior at the audited stage existed precisely because no production provider adapter was wired yet; the stub retained obligations instead of inventing evidence. That is admirable engineering and simultaneously a reminder: a verified mechanism in a remediation stage is not a shipped product capability. Product availability claims need product evidence — release records and verified access conditions — not architecture descriptions. Relatedly, target product boundaries (separate apps for specialized work, with one workspace spanning lightweight chat, cloud execution, and local environments) describe where responsibilities belong, not proof that every migration has shipped. Never convert a roadmap or a target separation into a present-tense availability claim.

Fourth, verifiability says nothing about task quality. A job can be perfectly verifiable and still do the wrong thing: match receipts to the wrong ledger, file on time but miscategorized, or optimize a metric nobody wanted. The checklist ensures the work can be inspected and its completion trusted; judging whether the work was worth doing remains a human responsibility, supported by review of the evidence the system preserved.

The one-page checklist

Copy this into your vendor review, architecture review, or build plan:

Outcomes. Explicit lifecycle states and transitions; one-way terminal states with stale writes rejected; executor completion separated from certified success; honest projection onto display labels.

Intervention. Approval, pause, block, cancel, and fail reachable from any unfinished state; approvals bound to specific actions with expiry; intervention durably recorded; authority re-checked at every dispatch.

Recovery. Unknown distinct from failed; durable leased reconciliation over committed obligations; mandatory re-observation after disruption; unknowns retained rather than invented; bounded replanning.

Evidence. Independent verifier with distinct privileges; coverage, freshness, and contradiction checks; non-empty evidence references enforced structurally; reconstructable history where summaries never override records.

A system that satisfies all four does not guarantee successful agents. It guarantees something more durable: when a job claims success, you can check — and when it cannot prove success, it says so plainly instead of asking for your trust.