Skip to content

EthenEthenEthen

When an Agent Action's Outcome Is Unknown

An agent action is interrupted mid-flight: did it happen? Ethen's mission reconciler refuses to guess — it retains the effect as unknown until evidence resolves it.

An agent action is interrupted mid-flight: did it happen? Ethen's mission reconciler refuses to guess — it retains the effect as unknown until evidence resolves it.

A mission agent dispatches an action. The request leaves the worker, the response never comes back, and the process restarts. At this point there are two tempting answers, and both are wrong: retrying blindly can execute the effect twice, while marking the attempt failed can abandon work that actually succeeded. The inspected Ethen missions reconciler — services/missions-worker/src/reconciler.ts — takes a third path. It records the attempt's effect as unknown, keeps the reconciliation obligation open, and resolves it only when an observation supplies evidence. This article explains how that loop works, why retention is the safe default, and what the Stage-0 exit audit does and does not establish about it.

Nothing here describes a live provider integration. The inspected reconciler injects its lookup providers, and the only live implementation is a test fake: there is no production provider adapter yet, so the wired serve loop observes unknown and retains every obligation rather than inventing evidence. Every concrete scenario below is labeled synthetic for exactly that reason.

Unknown is a state, not an error

Most retry logic works with two outcomes: the action succeeded or it failed. The missions design adds a third, and the independent Stage-0 exit audit of 2026-09-15 names it an explicit invariant: UNKNOWN is not FAILED. An effect flagged effect_known=false sits behind fenced indeterminate transitions in the action ledger, and the audit's phrasing is blunt — no timeout ever flips it. Only the absence of a lease holder resolves an unknown toward failed, and mission finalization rejects any mission that still carries an UNKNOWN_EFFECT.

That strictness exists because the two failure modes of guessing are both expensive. Treating an unknown as failed invites a retry of something that may already have happened: a duplicate charge, a second provisioning call, a repeated external write. Treating it as succeeded invites the opposite error: downstream steps build on an effect that never landed. Fencing the unknown — holding it in a state that neither path can silently exit — forces the system to go and look instead of assuming.

The fence is enforced where it matters most: completion. The audit records that the finalize path requires the mission to be VERIFYING with a verifier PASS, matching hashes, fresh observations, no unknowns, and every mandatory task SUCCEEDED. A task itself reaches SUCCEEDED only with a check reference and evidence references. Unknown effects therefore cannot leak into a completed mission by accident; they block the gate until they are resolved with evidence or the mission fails honestly for another reason.

Where unknowns come from

Unknowns are born at the boundary between the agent and the world it acts on. The reconciler's own code paths name the three ways an observation can come back empty-handed. First, the provider lookup can throw — the effect store is unreachable, the connector errors, the read times out. Second, the lookup can report a conflict, meaning the evidence disagrees with itself and no authoritative answer is available. Third, in the inspected Stage-0 wiring, the lookup can simply be the honest stub: no production adapter exists, so every observation returns unknown by construction.

Interruptions turn these transient conditions into durable obligations. A worker crash between dispatch and receipt, a restart that wipes process memory, a lease that expires while its holder is gone — each leaves an attempt whose effect may or may not have landed. The reconciler is built stateless across restarts for exactly this reason: all truth comes from committed rows, never from process memory, so a fresh worker can pick up the same obligations its crashed predecessor left behind. The obligation survives the interruption; the guess does not get made.

It is worth separating timeouts from knowledge here, because the design is deliberate about it. A timeout tells you that no response arrived in time. It says nothing about whether the effect landed on the other side — the request may have been processed a millisecond after the deadline. That is why the audit stresses that no timeout ever flips an unknown to failed. Time passing is not evidence; only an observation of the effect itself counts.

Claim, observe, resolve: the reconciler's loop

The reconciler is a serve-loop driver over reconciliation obligations, implemented as reconcileOnce. Each sweep follows the same three phases — claim, observe, resolve — and returns a summary of { claimed, resolved, unknown, errors } so operators can see what the sweep did. The defaults are a batch of 10 obligations and a 120-second lease, both overridable per call.

The sweep starts by listing obligations through the missions_list_reconciliation_obligations stored procedure: expired or open attempts that still need resolution. If even the listing fails, the sweep returns immediately with a single error rather than pretending it looked. Each obligation carries its identity and scope — attempt, tenant, project, environment, action intent, generation, attempt number, lease owner, and the stable operation key that names the effect.

Claiming comes next. For each obligation, the worker calls missions_claim_reconciliation with the tenant, project, environment, action intent, generation, its own worker ID, and the lease duration. The scope matters: the lease is bound to that exact tenant-project-environment slice, so two workers serving different scopes never fight over the same row, and two workers in the same scope resolve the race in the database, not in memory. Losing the race is routine, not failure — if the claim fails with MISSIONS_STALE_LEASE or MISSIONS_COMMAND_CONFLICT, the worker skips the obligation silently and moves on. Any other claim failure, or a claim that returns no attempt ID, is counted as an error and the sweep continues with the next row. One bad row never kills the sweep.

Observation is a single provider call keyed by the stable operation key: provider.lookup(operation_key). The lookup answers one question — is there a recorded effect for this key? — and returns whether it found one, optionally with a receipt and a conflict flag. The key is stable across retries and restarts, which is what makes the observation idempotent: asking twice about the same key asks about the same effect, never about two different attempts that happen to look alike.

Resolution writes the answer back through stored procedures. If the lookup threw or reported a conflict, the worker calls missions_resolve_reconciliation with the verdict unknown, and the obligation stays open for a future sweep. Otherwise it calls missions_resolve_reconciliation_effect with succeeded when the effect was found and failed when it was authoritatively absent. Only this last path — found or authoritatively missing — counts as resolved. Everything else is retained or counted as an error, and the sweep moves on.

Why retention is the default

Retaining unknowns is not a shelved error; it is the mechanism working as designed. The header comment in the reconciler states the Stage-0 boundary plainly: lookup providers are injected, the only live implementation is the test fake, and there is no production provider adapter yet — so the wired serve loop observes unknown, retaining every obligation, rather than inventing evidence. A reconciler that guessed would resolve faster and mean less. This one prefers an honest backlog over fabricated certainty.

Retention composes with the completion gate described earlier. Because finalization requires no unknowns alongside a verifier PASS, retained obligations cannot be papered over by a mission that otherwise looks done. They sit visibly in the unknown count of each sweep summary until a lookup with real evidence resolves them. The audit calls this an honest stub behind a typed seam: the EffectProvider interface is real, the crash behavior around it is tested, and the per-connector live adapters that would plug into it are explicitly deferred to later gates rather than faked.

There is a second, quieter reason retention is safe: the rest of the system re-observes rather than trusting stale state. The audit documents staleness fencing across the mission lifecycle — stale observations, stale world generations, stale revisions and contract generations are rejected at claim, task, and command boundaries, and the workflow layer re-observes on resume, wake, and dispatch. A retained unknown therefore re-enters a system that will check the world again before acting on it, instead of replaying a cached belief about what happened. Retention plus re-observation is what keeps an old interruption from corrupting a new decision.

Leases and exactly-once observation

Two mechanisms keep the loop itself from creating the duplicates it exists to prevent: scoped leases and stable keys.

The lease answers "who is allowed to resolve this obligation right now." Claiming binds the attempt to one worker for a bounded period — 120 seconds by default — inside one tenant-project-environment scope. Expiry returns the obligation to the pool instead of leaving it wedged on a dead worker, which is how the system survives crashes without manual intervention: the next sweep simply claims what the dead holder abandoned. And because lease races resolve as silent skips, running multiple workers for throughput does not multiply resolutions. At most one worker holds an obligation at a time, and losers move on without noise.

The stable operation key answers "which effect are we asking about." The audit describes mission-scoped idempotency keys, immutable receipts, post-action re-observation, and linked compensation as the Track C mechanism set, with generation fencing so that a repaired newer generation stales the older one rather than merging with it. Observing "exactly once by its stable key" — the reconciler's own description — means the lookup question is well-posed no matter how many sweeps run: the key names one effect, the receipt for it is immutable once written, and repeated observations converge on the same answer instead of minting new effects.

Recovery behavior around restarts reinforces both. The audit's Track H findings record committed-rows-only restore with UNKNOWN preservation: control state is reconstructed from the database, unknown effects survive checkpoint restore, and recovery proceeds within fixed bounds rather than retrying forever. A worker can die mid-sweep and its replacement sees the same obligations, the same keys, and the same unknowns — continuity comes from the rows, not from the process.

A synthetic walkthrough

The following example is synthetic: it illustrates the inspected mechanism with an invented action, since no live provider lookups exist in the inspected code.

Imagine a mission step whose action is "provision a review virtual machine" with operation key op_synth_042. The worker dispatches the request, the provider starts the VM, and the worker crashes before the receipt is recorded. On restart, the action ledger holds an attempt with an unresolved effect — unknown, fenced, blocking nothing yet but completing nothing either.

The next reconciler sweep lists the obligation and claims it with the new worker's ID and a fresh 120-second lease. It then calls provider.lookup("op_synth_042"). In the inspected Stage-0 wiring there is no production adapter, so this observation comes back unknown, and the worker resolves the sweep entry as unknown. The obligation is retained. The sweep summary shows one claimed, one unknown, zero resolved — an honest report that the question is still open.

Now consider the design-target half of the walkthrough, which is also synthetic and describes the seam's intended use rather than inspected behavior. Once a certified per-connector adapter is plugged into the injected provider, a later sweep claims the same obligation and performs the same lookup against the provider's effect record. If the adapter finds the VM's creation receipt under op_synth_042, the worker resolves the effect as succeeded exactly once; if the provider authoritatively reports no such VM was ever created, it resolves failed. Either way, one observation of one stable key settles the question every earlier sweep had retained.

The walkthrough shows why the operation key must be stable and mission-scoped. Had the retry minted a fresh key, the lookup would have asked about a different effect than the one the crashed worker dispatched — and a second VM could have been provisioned while the first ran unnoticed. Same key, same question, one answer.

What the audit established — and its limits

The claims above rest on two inspected sources: the reconciler implementation and the independent Stage-0 exit audit dated 2026-09-15. The audit is a dated report, so it proves only its own scope — but within that scope it is specific and independently corroborated.

On the mechanism side, the audit records Track C (reconciliation) as PARTIAL in a precise sense: the reconciliation mechanism plus the crash matrix are green, while per-connector live adapters are absent by design. The crash matrix covers cases M1 through M11 — timeout collapsing to unknown and then reconciling to success or failure exactly once, invoke-once semantics, exactly-once settlement and release, marker races, and a backlog proof past 500 entries. Track H (reconstruction and recovery) is a full PASS, including UNKNOWN preservation across checkpoint restore and restart survival. The technical gates re-run green in the audit window: root typecheck and build pass, the behavioral suites pass, missions unit and integration suites pass at 230 of 230 each, and the certifier passes 7 of 7 stages.

On the boundary side, the audit is equally explicit about what is deferred. Live provider reconciliation sits outside the Stage-0 contract itself, classified with later gates G4 through G10; the exit verdict authorizes Stage-1 development only, with sandbox, egress, provider certification, and production gates still closed. The gap register lists live per-connector effect reconcilers as a material later-gate item owned by post-Stage-0 provider certification, naming the injected EffectProvider as the seam. Nothing in the audit supports claims about live provider behavior, production availability, or certified autonomy tiers for missions — the audited object is the mechanism, tested with fakes and deterministic suites, not a operated production loop.

That scoping is a feature of the evidence, not a gap in it. It tells readers exactly which half of the story is implemented and tested (claim, fence, retain, resolve-by-evidence) and which half is specified interface awaiting adapters (live effect lookups per connector). This article stays on the implemented side throughout; the synthetic walkthrough above is labeled wherever it steps toward the seam's intended use.

Keep the question open until evidence closes it

The reconciler's lesson generalizes beyond missions infrastructure. Whenever an agent acts on the world through an unreliable boundary, "did it happen" is a question for observation, not for timeouts or optimism. The inspected design answers it with a small, strict loop: claim the open question under a scoped lease, ask about one stable key, write down exactly what the evidence says — and when there is no evidence, keep the question open. Retention looks like inaction. It is actually the system refusing to trade an unknown it can see for a certainty it cannot prove.