Inside Ethen's Durable Job Service
Leases decide who may act, envelopes decide who the work belongs to, and reconciliation decides what actually happened.
Leases decide who may act, envelopes decide who the work belongs to, and reconciliation decides what actually happened.
AI work keeps running after the process that started it is gone. A research job fans out to a provider, the worker crashes mid-call, and a second worker picks up the same row minutes later. Did the provider call land? Is it safe to send it again? Ethen durable AI jobs answer those questions in one shared implementation, the DurableJobService, instead of leaving every product to invent its own recovery story. This article traces what that service actually does: how jobs are created exactly once, how leases with fencing generations keep stale workers from mutating state, how execution envelopes bind each job to a tenant and actor, and how reconciliation resolves uncertain outcomes without blind retries.
One service, one job record
The service centers on a single JobRecord. Every job carries an organization, a project, a queue name, a payload, an idempotency key, a status, priority and attempt bookkeeping, lease fields, cancellation fields, scheduling fields, dispatch-boundary fields, escalation fields, and reconciliation metadata. Creation starts every job in queued with attemptCount at zero, no lease, and no error. The defaults are small and explicit: queue "default", three maximum attempts, a 30-second backoff base, and a 60-second lease duration.
Creation is idempotent by key. createJob requires a non-empty organization, project, idempotency key, and payload, then delegates to the repository, which returns the existing job when the same project-scoped idempotency key already exists. Retried submissions therefore cannot silently mint duplicate work; the caller can also look the job up afterward with findByIdempotencyKey to confirm suppression from its own side. Scheduling is a field, not a separate system: scheduledAt defaults to now, and the claim path only considers jobs whose scheduled time has arrived.
The status model is a closed set of ten states: queued, claimed, running, completed, failed, cancelled, dead_letter, indeterminate, halt_unsafe, and timed_out. Seven of those are terminal. The transition table is deliberately narrow. A queued job can only be claimed or cancelled. A claimed job can move to running, return to queued, fail, cancel, dead-letter, or enter one of the three uncertain or halted states. A running job can complete or reach the same failure-family outcomes. Terminal states have no onward transitions in the table at all. Two states carry extra rules worth remembering now: indeterminate resolves only through the explicit reconcile-dispatched path, and halt_unsafe never transitions onward under any path.
That narrowness is the point. Every state change a job can undergo is enumerable from one table, which means recovery logic never has to guess what a status means. When a worker disappears, the next worker reads the row and the table together and knows exactly which moves are still legal.
Leases: claiming work, proving you still own it
No worker executes a job it has not claimed. claimJob atomically takes the next eligible job for a worker id, with an optional organization scope so a worker can be restricted to one tenant's work. Eligibility has four conditions: the job must be queued, or be claimed or running with an expired lease; its scheduled time must have arrived; and it must not carry a pending cancellation request. Among eligible jobs the ordering is deterministic — priority descending, then scheduled time ascending, then creation time ascending — so two workers racing for work converge on the same candidate and the atomic claim decides the winner.
Claiming does three things at once. It flips the job to claimed, increments the attempt count, and stamps the lease: the worker id becomes both leaseId and claimedBy, an expiry is computed from the lease duration, and — critically — the lease generation is incremented. That generation is the fencing mechanism. It starts at zero when the job is created, and every successful claim, including a reclaim after a lease expires, mints a strictly newer number. The comments in the code state the reason plainly: worker ids may be reused, so worker ids alone cannot prove freshness. Only the generation can.
Once a worker holds a claim, it must keep proving it. renewLease and its alias-friendly entry point heartbeat extend the lease and promote a claimed job to running, but they refuse in three situations: the caller does not own the lease, the presented generation does not match the stored one, or cancellation has been requested. A stale worker — one whose lease expired and whose job was reclaimed by someone else — presents an old generation and gets a false. Its subsequent mutations fail the same fencing check. The design assumes crashes and slow workers as the normal case: expiry plus reclaim keeps work moving, while fencing guarantees the expired holder cannot corrupt the job after losing it.
Consider the concrete failure this prevents. A worker claims a job at generation 4, stalls in a long provider call past its 60-second lease, and the job is reclaimed at generation 5 by a healthy worker. The stalled worker finally wakes up and tries to record completion. Without fencing, its write would land on a job it no longer owns, possibly overwriting the new holder's dispatch marker. With fencing, the generation mismatch rejects the write. The stale worker learns it lost the lease — the Research worker surfaces exactly this as a LEASE_LOST outcome — and the job's truth stays with the current generation.
Execution envelopes: identity that cannot be forged
The payload of a job carries the work to do, but identity travels in reserved envelopes inside that payload. There are two: the execution envelope, holding canonical execution identity, and the governance envelope, holding queue-time governed dispatch facts. Both can only be written through typed creation inputs. If a raw payload carries an execution key without a typed execution input, or a governance key without a typed governance input, creation is rejected. Tenant, actor, and governance bindings cannot be smuggled in through untyped fields — the code calls this out as an anti-forgery rule.
The execution envelope is built by buildExecutionEnvelope from two sources that the caller cannot confuse: the typed input supplies tenant, actor, trace, and lineage identifiers, while the authoritative job record supplies organization, project, and the job id itself. The resulting identity reads as a lineage chain — tenant, organization, project, actor, task, run, attempt, job, and trace — plus links to the action intent, admission decision, and approval when governed execution is in play. Tenant, actor, and trace are required; a consequential dispatch without them fails closed.
Reading is as strict as writing. readExecutionIdentity throws on a malformed envelope rather than returning partial identity, and validateExecutionIdentityForDispatch performs a fail-closed readiness check before dispatch: identity must exist, tenant, actor, and trace must be present, and the organization, project, and job recorded in the envelope must match the job row. Any mismatch produces a reason, and any reason means no dispatch. The governance envelope follows the same discipline. It carries the action intent plus the capability, policy, and budget facts the queue-time authority established; the worker re-verifies digest integrity and expiry at dispatch time, and approval revocation and consumption are re-verified live against the canonical approval store rather than trusted from the envelope alone.
The effect is that identity is never reconstructed from scattered payload fields. There is one canonical place to look, one function that writes it, one function that reads it, and a validator that refuses to proceed when the envelope and the row disagree. For a shared service that may one day schedule work across products, that strictness matters more than convenience: a forged or drifted identity could otherwise attach one tenant's side effects to another tenant's ledger.
The dispatch boundary: marking the point of no blind return
The most consequential single write in the service is markProviderDispatched. The instant a worker dispatches a provider or external side effect, it records the dispatch boundary: a timestamp plus a stable reconciliation operation key. The call is owner-only, fenced by generation like every other mutation, and it promotes a claimed job to running. From that moment, the job's failure semantics change completely.
Before the boundary, a failure is clean. The worker can fail the job as retryable, and if attempts remain the job returns to queued with exponential backoff — base seconds doubled per attempt, capped at one hour — for another worker to claim. After the boundary, a failure is uncertain. The provider may or may not have acted, so a blind retry could duplicate a real-world effect. The contract is explicit: once the dispatch boundary is recorded, a retryable failure must not blind-retry. It lands in indeterminate instead, and indeterminate can only be resolved by reconciliation against the stable operation marker.
Three owner-only methods terminalize claimed or running jobs with a required error record. markIndeterminate records the uncertain outcome. markHaltUnsafe records a state from which the job must never move again — the code provides it for situations where onward transition itself would be dangerous. markTimedOut records a bounded worker execution that ran out of time. Each requires an error code and message, so the terminal row always explains itself.
The Research consumer shows the boundary discipline in practice. Its handler heartbeats before dispatch, then marks the provider boundary with a stable key derived from the job before any external provider effect: research:<jobId>:<mode> for research lanes and deep-research:<jobId>:<runId> for deep research. The durable worker repeats the pattern — cancellation check first, heartbeat owner check, dispatch-boundary write, then the provider call — and on error after the boundary it writes indeterminate rather than failing cleanly. Losing the lease surfaces as LEASE_LOST, the one outcome the handler treats as retryable, because a lost lease means another generation now owns the truth.
Reconciliation: resolving uncertainty by evidence, never by guessing
Uncertain jobs leave the candidate set only through reconciliation or operator action. findReconciliationCandidates returns the project-scoped sweep input: jobs that are indeterminate or timed_out at any age, plus claimed or running jobs with expired leases. The query never crosses project scope and orders oldest first, so sweeps are deterministic — the same store state always produces the same outcome.
The sweeper itself, sweepReconciliationCandidates, is a control loop with a conservative temperament. For each candidate it re-checks the project binding, then applies a series of skips and resolutions. Already-escalated jobs are skipped until an operator acts. Crashed-but-undispatched jobs — claimed or running with an expired lease and no dispatch marker — are left for the claim and reclaim path; the sweeper inspects but never seizes live work. Timed-out jobs are escalated, because a timeout requires external verification before any resolution. Indeterminate jobs without a stable operation marker are escalated too, since there is nothing to verify against.
Only one case resolves automatically: an indeterminate job with a stable operation key. The sweeper queries authoritative downstream state by that key. If the provider reports success, the job reconciles to completed. If it reports failure, the job reconciles to failed, or to dead_letter when the deployer configures policy failures that way. If the provider state is unknown, the job is escalated with a reason naming the operation key, and automatic resolution stops. The sweeper never re-dispatches work just because the status is uncertain. Every resolution and escalation appends to the job's durable event trail, so the recovery itself leaves evidence.
That event trail deserves a final note. The service exposes appendJobEvent and listJobEvents: an append-only per-job history over a closed vocabulary of lifecycle events — claimed, dispatched, heartbeat, completed, failed, cancelled, indeterminate, halt-unsafe, dead-lettered, timed out, retried, admission-checked, escalated, reconciled, and an execution-control shadow event. Workers record claim, dispatch, and terminal transitions so operators can reconstruct what happened across crashes and reclamations. The vocabulary is enforced: the Research handler's comments note that custom event names are rejected by the durable event writer, and that such a rejection once masked a real terminal state as a persistence error. Completion is recorded by the worker host as the canonical completed event carrying the result — handlers do not invent their own terminal vocabulary.
Cancellation, retries, and dead letters
Cancellation respects the lease. Cancelling a queued job transitions it to cancelled immediately. Cancelling a claimed or running job only plants a flag — a timestamp and reason — that the owning worker must acknowledge with acknowledgeCancellation, which performs the terminal transition. While the flag is set, lease renewal and heartbeats return false, so a worker that checks its heartbeat learns promptly that it should stop. The Research durable worker checks for cancellation before dispatch and acknowledges it directly, returning success without touching the provider.
Failure handling follows the attempt budget. failJob takes an explicit retryable flag: retryable with attempts remaining reschedules with backoff, exhausted retries move the job to dead_letter with a timestamp and reason, and non-retryable errors move it to failed. Operator recovery is deliberately asymmetric. retryJob re-enqueues only failed jobs. Dead-lettered jobs stay terminal — the dead letter is the durable record that the budget ran out — and indeterminate jobs are not retryable at all, because only reconciliation on the stable marker can resolve them. Uncertainty is never silently converted into a retryable failure, and the escalation path keeps timed-out and indeterminate jobs visible to operators rather than letting the sweeper guess.
What this traces — and what it does not prove
This article traces the implemented shared service: the DurableJobService methods, the contract table, the envelope builders, the repository interface, the deterministic sweeper, and one real consumer in the Research worker handler. A few boundaries of that evidence should stay explicit.
First, describing a shared service is not a claim that every product persists work identically. The Research handler's own comments say there is no second job platform — research consumes the single durable-jobs service — but that statement covers Research. Other target apps (Studio, Designer, Founder, and the rest) are separate products with their own code paths, and this article asserts nothing about which persistence each one uses. Target separation is an architecture fact, not proof of any migration.
Second, the code describes worker behavior, not certified live workers. The lease durations, fencing checks, heartbeat discipline, and LEASE_LOST handling are implemented logic; the inspected sources do not certify that production workers are running, how many exist, or what throughput they sustain. No claim here should be read as an availability or performance statement.
Third, some pieces are interfaces their deployers must fill. The sweeper's provider-state query is a supplied function: reconciliation is only as authoritative as the downstream lookup the deployment wires in. The repository interface has multiple possible backends — the inspected tree includes both in-memory and Supabase-backed implementations — and behavior verified against one backend is not automatically proven for another. Timed-out and indeterminate jobs ultimately depend on operators and external verification, by design.
What the implementation does establish is a coherent durability contract in one place: idempotent creation, fenced leases, unforgable identity, a marked dispatch boundary, evidence-driven reconciliation, and an audit trail over a closed event vocabulary. For Ethen durable AI jobs, surviving the crash is not a feature of any single product. It is the shared service's entire job.