Making Mission Completion Depend on Evidence
In the Stage-0 mission system, running the work and proving the work are different transitions — and only independent evidence unlocks the second.
In the Stage-0 mission system, running the work and proving the work are different transitions — and only independent evidence unlocks the second.
The executor does not get the last word. In the Stage-0 mission system, the component that performs a task cannot mark that task successful on its own. A separate verification step, backed by evidence references and an independent check, stands between "the work ran" and "the work counts." That separation is the working definition of Ethen mission completion evidence: completion is a decision the system makes about evidence, not a status a worker claims about itself.
This article explains how that decision works at the code level: the mission and task state machines, the pure reducers that guard each transition, the constraint ledger that narrows authority before dispatch, and the SQL finalization gate that holds the actual authority to complete a mission. It also states the boundary plainly. The Stage-0 suites verify the mechanism under deterministic conditions. They do not certify real-world autonomy, live tool use, or unattended operation.
Two state machines, one rule
The domain state machine defines 16 mission states and 9 task states, each with an explicit transition table. The mission list runs from early contracting through execution and verification to a terminal outcome:
CREATED, CONTRACTING, PLANNING, READY, RUNNING, WAITING, AWAITING_APPROVAL, BLOCKED, RECONCILING, RECOVERING, VERIFYING, COMPLETED, FAILED, ABORTED, PAUSED, and CANCELING.
The task list is smaller and sits underneath:
PENDING, READY, RUNNING, WAITING, BLOCKED, VERIFYING, SUCCEEDED, FAILED, and CANCELED.
Three mission states are terminal — COMPLETED, FAILED, ABORTED — and three task states are terminal — SUCCEEDED, FAILED, CANCELED. The terminal rule is absolute: the transition function for a terminal state returns no outgoing edges, and any attempt to apply a transition out of one is rejected with a terminal-state error. A mission that reaches COMPLETED cannot be reopened by a later write, and a task that reaches SUCCEEDED cannot be recycled into a new run. If the world changes after that point, the system must start new work rather than rewrite the recorded outcome.
The happy-path edges read like a pipeline with deliberately narrow doors. A mission moves CREATED to CONTRACTING to PLANNING to READY to RUNNING, and a running mission can step to WAITING and back, or forward into VERIFYING. From VERIFYING the mission has exactly three exits: COMPLETED, RECOVERING, or FAILED. There is no direct edge from RUNNING to COMPLETED. The same shape holds for tasks: RUNNING can move to VERIFYING, and VERIFYING can move to SUCCEEDED, back to READY for another attempt, or to FAILED. Skipping the verification state is not a supported transition at either level.
Interruptions are handled by a second edge family rather than by ad-hoc jumps. Approval requests, blocks, pauses, cancellation, and fatal failure can interrupt any active mission state: from almost anywhere, the mission may step to AWAITING_APPROVAL, BLOCKED, PAUSED, CANCELING, or FAILED. Three restrictions keep this family disciplined. It never offers a self-loop, it never applies from inside CANCELING, and it never applies out of a terminal state. Cancellation itself carries a closed set of exits — RECONCILING, ABORTED, or BLOCKED — so that a mission being wound down can quiesce, abort, or park on unresolved effects, but cannot approve new work, pause itself again, or relabel failure from inside the cancellation path.
Pure reducers propose; they do not dispose
The most load-bearing sentence in the state-machine module header is a negative claim: the reducers are pure. Time, identifiers, and evidence references arrive through function arguments. The transition functions perform no input or output, draw no randomness, and read no wall clock. Given the same current state, target state, and context, they return the same verdict every time.
Each guarded transition is a compare-and-swap over a revision number. The caller passes the version it believes is current; if the stored revision differs, the transition is rejected as stale and the write fails closed. Stale here covers the realistic concurrency hazards: two writers racing on the same mission, a resumed worker acting on a snapshot taken before a human intervened, a retry replaying an old command. The stale writer loses, and nothing is half-applied. Unknown states, illegal edges, and terminal writes are rejected with their own error codes, so a caller can distinguish "you asked for a transition that does not exist" from "someone else moved this record first."
This purity is a strength and a limit at once, and the distinction matters for everything that follows. The reducers decide whether a requested transition is legal. They do not decide whether a mission is complete in any durable, authoritative sense. They hold no database connection, enforce no cross-record invariant, and see no verifier verdict beyond the reference string the caller supplies. The September 15 independent Stage-0 exit audit puts the production relationship bluntly: the TypeScript admission path is advisory-only, and the database decides. Model-side code proposes actions and transitions; the SQL layer authorizes dispatch and finalization. Any reading of the reducers as the completion authority mistakes the guard at the door for the court inside.
A small interop detail reinforces the point. The module ships a projection from the 16 mission states onto a lowercase run vocabulary, documented explicitly as an approximation for interop only — never a claim that a mission is a run. Unmapped future states surface as errors rather than collapsing into a near neighbor.
The task path to SUCCEEDED runs through VERIFYING
Task success carries one extra guard that mission transitions do not: entering SUCCEEDED requires an independent-check reference. The transition context for tasks includes an optional check reference field, but it is optional only in the type. When the target state is SUCCEEDED and that field is missing or empty, the reducer rejects the transition with a missing-check error. The header comment states the rule without hedging: executor completion alone never unlocks dependents.
Submission for verification is itself a guarded step with its own evidence requirement. The helper that moves a task from RUNNING to VERIFYING takes a list of evidence references and rejects an empty or missing list before it ever reaches the transition table. Its documentation adds the second half of the rule: submitting evidence grants no certification privilege. Moving into VERIFYING is a request for judgment, not a judgment. SUCCEEDED still needs the independent check, and that check comes from outside the executor.
The resulting task lifecycle has a shape worth tracing. A task waits in PENDING, becomes READY, and starts RUNNING. From there it may wait on a dependency, block, fail, be canceled, or submit evidence and step into VERIFYING. Inside VERIFYING there are three exits and only three: SUCCEEDED with a valid check reference, back to READY for rework, or FAILED when the work is judged unrecoverable. The loop between READY, RUNNING, and VERIFYING is where retries live, and nothing in that loop can mint success without passing through the checker.
The mission path to COMPLETED runs through SQL
Mission completion follows the same two-door pattern at a larger scope, except the second door is a database function, not a TypeScript reducer. According to the September 15 independent exit audit, the finalization routine requires the mission to be in VERIFYING, plus a verifier-side PASS verdict, plus evidence-hash and freshness checks, plus a no-unknown-effects condition, plus mandatory tasks in SUCCEEDED. Every conjunct must hold. A mission with passing checks but one unresolved effect, or with complete tasks but stale evidence, does not complete.
This is where the pure-reducer versus SQL-authority distinction becomes concrete. The TypeScript layer can walk a mission from RUNNING to VERIFYING when asked legally, and it can reject illegal requests deterministically. But the step from VERIFYING to COMPLETED in production is owned by the finalization gate, which re-reads the facts that matter — revisions, contract digests, world generations, approvals, budgets, kill-switch state — at claim time rather than trusting what any caller asserts. Stale bindings, revoked grants, and drifted digests deny the claim instead of degrading gracefully.
The audit's invariant table gives this design its slogan: the model proposes, and the system authorizes. Proposal stages intent; only the system's own re-reads and checks convert intent into effect. Completion is the strictest instance of that pattern, because it is the transition the outside world will rely on. Everything the reducer checks cheaply in memory, the finalizer re-checks authoritatively against durable state before COMPLETED becomes real.
Two corollaries follow. First, unit tests over the reducers prove transition logic, not completion authority; finalization tests must run against the SQL gate with its verifier, hash, freshness, unknown-effect, and task-coverage conditions. Second, any future change to what completion means — a new evidence type, a tighter freshness bound, an additional mandatory check — must land in the finalizer to be authoritative. Changing the TypeScript table alone would move the guard rails without moving the gate.
The verifier stands outside the executor
Independent verification earns the word independent through deployment and privilege separation, not merely through a separate function name. The exit audit records the verifier as a separate service with its own process, container image, role, and queue. The worker and verifier connect under restricted database roles rather than a shared superuser, and the audit's grep over all 39 files in the worker, verifier, and missions package found zero service-client constructors — the privileged bypass path does not exist in that code. Finalization and result-recording grants belong to the verifier role, and executor-side attempts to finalize are denied and covered by tests.
What the verifier checks is deliberately unglamorous. It re-hashes evidence bytes rather than trusting a claimed digest, applies deterministic oracles rather than model judgment, and enforces a 24-hour freshness window with a 5-minute clock-skew allowance, so old observations cannot quietly authorize new completions. Coverage checks catch thin evidence sets, contradiction registers block on high-severity conflicts, and the audit records false-green and false-red test matrices over the verifier with appeals handled as their own path.
The headline consequence, stated as an invariant rather than an aspiration, is that reaching COMPLETED without independent evidence is impossible through the finalization path: the gate derives its verdict from the verifier's PASS over the current contract, rubric, and plan, combined with hash, freshness, and mandatory-task conditions. An executor can do everything right and still sit in VERIFYING until the verifier agrees. That asymmetry is intentional. It converts "trust the worker" into "inspect the evidence," and it gives reviewers a stable artifact — the verification record — to audit after the fact.
The constraint ledger narrows authority before anything runs
Completion gating answers "may this finish," but an earlier question gates every action along the way: "may this run at all under this mission's contract." The constraint ledger answers that question, and its governing rule is stated in the first lines of the module: mission constraints only narrow Platform authority, never widen it.
The ledger compiles seven deterministic hard-predicate kinds into pure evaluator functions: environment equality, domain allowlist, no-external-messaging, maximum spend in micros, approved-dataset-only, no-source-deletion, and no-production-writes. Each predicate carries a key, a kind, and parameters, and each compiled check takes a proposed action and returns either null for allow or a human-readable denial string. A spend check compares digit-only micro-amount strings as big integers so that arithmetic never drifts through floating point. A dataset check denies undeclared datasets outright rather than treating silence as consent. A production-write check fires only when both the write flag and the production environment coincide, keeping staging writes usable under the same contract.
Three fail-closed behaviors make the ledger trustworthy as a gate. First, unknown or malformed predicates fail compilation instead of passing: an unsupported kind, a missing allowlist, or a non-numeric spend limit becomes a denial reason, and contracting blocks rather than degrading to model judgment. Second, manifest evaluation denies on any violation, including uncompilable entries — one bad predicate poisons the action, and there is no partial-credit path where six passing checks outweigh one failing check. Third, manifests are cached and trusted by content digest, never by mutable pointer: the same digest hits the cache, while any amendment produces a new digest and recompiles. A missing manifest, a digest mismatch, or a concurrent amendment denies dispatch.
The audit connects this ledger to the completion story through invalidation. Amending a contract invalidates outstanding approvals, proposals, and certification artifacts keyed to the old digest, and the claim path re-checks the live grant, kill switch, and ceiling on every attempt. Authority can shrink mid-mission — a tighter spend cap, a revoked grant, a flipped kill switch — and the system enforces the shrinkage at the next gate rather than honoring a stale permission. That is what "shrinks, never silently expands" means in operation: every narrowing takes effect, and no widening happens except through an explicit, re-verified new agreement.
UNKNOWN is a fenced state, not a quiet failure
Distributed work produces a third outcome beyond success and failure: the action was attempted, but its effect is unknown. A timeout fired before the receipt arrived. A worker crashed between dispatch and observation. A provider acknowledged ambiguously. The Stage-0 design treats this outcome as a first-class fenced state rather than rounding it to either pole.
At the ledger level, the action record carries an explicit known-or-unknown flag on its effect, and the audit quotes the governing rule: no timeout ever flips it. Timeouts produce UNKNOWN, and only a defined resolution path — re-observation by the lease holder, reconciliation against provider truth, or lease-holder absence resolving to failure — moves the record out. The SQL transitions into and out of the indeterminate marking are fenced, and the finalization gate rejects missions with unknown effects still outstanding. A mission cannot complete while any consequential action's outcome remains unresolved.
The crash-matrix suites exercise this path systematically: invoke-once semantics, exactly-once settlement, marker races, and backlog draining past 500 entries, with timeouts resolving to unknown and then reconciling exactly once. Checkpoint-restore tests confirm UNKNOWN survives restarts, and verifier-finalization tests confirm it blocks completion even when everything else passes. UNKNOWN is allowed to exist, forbidden to hide, and barred from the COMPLETED door until resolved.
A worked example: one task, two gates
Consider a mission with a single mandatory task: fetch a dataset, transform it, and store the result. The mission walks CREATED to CONTRACTING to PLANNING to READY to RUNNING as its contract compiles, its plan forms, and execution begins. The task walks PENDING to READY to RUNNING as it becomes eligible and starts. So far every transition is a reducer check plus a revision bump, and any stale writer or illegal jump is rejected with a precise error.
The worker finishes the transform and wants credit. It cannot transition the task directly to SUCCEEDED — no such edge exists. Instead it calls the verification-submission helper with its evidence references: the dataset identifier, the transform parameters, the output hash, the storage receipt. If that list is empty, the call fails with a missing-check error before touching the state table. With references present, the task steps RUNNING to VERIFYING. This step records the claim and its supporting pointers. It certifies nothing.
The independent verifier now takes over: it re-hashes the output, checks the dataset against the approved list, confirms spend stayed under the micro-budget cap, verifies freshness, and scans for contradictions. Suppose the hash matches but the source observation aged out during a long transform. The verifier withholds PASS, and the task returns VERIFYING to READY for another attempt. The executor re-runs with a fresh observation, re-submits, and this time the verifier records PASS with a check reference.
Only now can the task step VERIFYING to SUCCEEDED, supplying that reference to satisfy the guard. And only with the mandatory task in SUCCEEDED, the mission in VERIFYING, the verifier's PASS current, the hashes matching, the evidence fresh, and no unknown effects outstanding can the SQL finalizer step the mission to COMPLETED. Remove any one conjunct — expire the evidence, lose the check reference, leave a compensation unreconciled — and the mission waits. Each failure mode names its own remedy: re-observe, re-verify, reconcile, then ask again.
What Stage-0 proved, and what it explicitly deferred
The September 15 independent exit audit is the right lens for calibrating claims, because it separates the contracted Stage-0 scope from everything deliberately postponed. Its verdict was PASS for Stage-0 exit with zero blocking issues, authorizing Stage-1 development only. Capability enablement — sandboxing, egress, provider certification, production operation — remained behind gates G4 through G10, which the audit records as still closed. A dated report proves only its scope, and this report's scope is the mechanism under test, not live autonomy.
Within that scope, the evidence is substantial. The audit re-ran six technical gates from the current tree: root typecheck and build green, the full behavioral suite green across hundreds of files and thousands of assertions, missions unit and integration suites green at 230 passing tests each with zero skips, and the seven-stage certifier green end to end. The deterministic MissionBench harness ran 15 of 15 cases over 33 trials — honestly labeled as pure checkers with no providers, network, or generated code involved.
Outside that scope, the audit's track scorecard is mostly PARTIAL by design. The browser-process boundary is implemented and tested, but no certified code-execution sandbox exists. Stale-UI rejection and takeover paths are tested, but no live page trials ran. The reconciliation mechanism and crash matrix are green, but no per-connector live adapters exist — the live-lookup seam throws by design rather than faking the answer. Supervised-only operation is certified; autonomy-tier promotion is explicitly excluded, and the audit confirms no autonomy-level enum exists in the missions code. Schedules, recurring missions, continuous operation, and the model-judge seam with its calibration dataset are all absent.
None of this is a quiet footnote. The audit classifies the deferred items as later-gate requirements with exact requirement names — certified sandbox adapter, provider reconcilers, live browser trials, live-model bench trials, judge plus calibration data, schedule engine — and states plainly that live sandbox, browser takeover, provider reconciliation, model judgment, continuous autonomy, and production concerns were excluded by the Stage-0 contract itself. Stage-0 tests therefore prove that the evidence gate holds under the conditions tested. They do not prove that an agent operating live tools, long horizons, or untested providers would complete real work safely. That proof would require the deferred gates, and the audit does not claim them.
Limits to carry forward
Even inside the passing scope, the audit leaves named work for Stage 1, and honest adopters of this design should track it. One item concerns guard wording: the verification-request path's READY-to-VERIFYING allowance sits against a blueprint that describes RUNNING-to-VERIFYING with all tasks submitted. The completion authority in the finalizer is strict regardless, so this is a confirm-or-tighten item rather than a hole — but it shows the audit doing its job, flagging even a wording-level mismatch between blueprint and gate. A second item notes the absence of a behavioral unit test for the verifier's world-observation oracle. A third proposes a monotonic authority-envelope version check at claim time as defense in depth, while recording that per-action live fencing already covers the expansion threat the check would address. Legacy unlinked rows keep permissive transitions by design, with safety holding for linked rows; that design choice deserves a re-read before any migration widens the linked set.
Human calibration gets its own explicit limit: the verifier mechanism passed, but measuring how well humans and the verifier agree in practice remains a later-gate requirement. Replanning is bounded by count plus generation fencing rather than by an error deadband. These are specific limitations owned by named gates — not blanket certification, and not construable as such.
The practical takeaway for teams borrowing this pattern is to preserve the layering. Keep the reducers pure so their logic stays exhaustively testable. Keep the finalizer authoritative so no in-memory shortcut can mint completion. Keep the verifier separate in deployment and privilege, not just in name. Keep UNKNOWN fenced and visible until reconciled. And keep the test claims scoped: deterministic suites prove the gate holds under the conditions they cover, and every condition they do not cover — live tools, live pages, live providers, live judgment — needs its own gate before anyone calls the system autonomous.
A mission that completes under these rules has earned its terminal state twice: once when the executor produced evidence worth judging, and once when an independent checker judged it sufficient. The state machine makes the two moments unmistakable, the ledger keeps the authority narrow throughout, and the SQL gate refuses to confuse the first moment with the second.