Skip to content

EthenEthenEthen

Separating Automation Triggers From Execution

Ethen's trigger boundary accepts an event only if it matches its subscription, records it once, and hands execution to someone else.

Ethen's trigger boundary accepts an event only if it matches its subscription, records it once, and hands execution to someone else.

Most automation bugs live at the seam between "something happened" and "do something about it." A webhook fires twice and the workflow runs twice. A poller restarts and replays a page it already processed. A renamed operation keeps receiving events authorized for its old identity. Each of these is a trigger problem, not an execution problem — and Ethen's automation trigger design treats that distinction as a hard boundary. The trigger layer validates, deduplicates, and enqueues. It never runs the workflow itself.

This article traces that boundary through one inspected source: the TriggerService implementation in lib/flows/triggers/trigger-service.ts. Every behavioral claim below comes from that file's 124 lines. Where the file ends, I say so — because the most important fact about this code is stated in its own comments: the inspected service is a process-local adapter for tests and local development, and production durability requires database backing that is not part of what I inspected.

The boundary in one sentence

The file's header comment states the design contract directly: provider calls and run execution are deliberately injected, so accepting an event can only enqueue a durable request. That single sentence explains the shape of everything that follows. The service knows about subscriptions, events, cursors, and schedules. It does not know how to call a provider or how to run a workflow — those capabilities arrive as interfaces passed in by the caller.

Two interfaces define the seam. RunEnqueuer exposes exactly one method, enqueue, which takes a tenant, a workflow identity and version, an event ID, and an idempotency key. There is no execute, no callback, no result channel. ProviderSubscriptionLifecycle exposes create, renew, and delete for provider-side subscriptions. The trigger service orchestrates when these are called, but the implementations live outside it. This is what "separating triggers from execution" means concretely: the trigger layer's only path to causing work is a narrow enqueue call carrying an idempotency key.

Subscriptions are fixed bindings

Every event in this design belongs to a subscription, and every subscription is a fixed binding between an external occurrence and an internal workflow. The TriggerSubscription record carries the full binding: a tenant ID, a workflow ID plus version, an operation ID plus version, a connection ID, and a trigger kind. It also carries lifecycle state — a status of active, paused, expired, or error — and kind-specific tracking: a provider subscription ID, a renewal timestamp, a poll cursor, an overlap window in milliseconds, and, for schedules, the next run time, a timezone, and a misfire policy.

New subscriptions start in a known state. createSubscription assigns a trg_ identifier, marks the subscription active, and leaves the provider subscription ID, renewal time, and cursor empty. Nothing about creation contacts a provider; provisioning is a separate explicit step. The provision method takes a subscription ID and a provider lifecycle implementation, calls create on the provider, and records the returned provider subscription ID and renewal timestamp. Creation and provisioning are different operations, which means a subscription can exist locally before — or without — a provider-side registration.

The fixed nature of the binding matters most when events arrive. The accept method rejects any event whose tenant, connection, operation, operation version, or trigger kind does not exactly match the subscription, or whose subscription is not active. The error message is precise: "Trigger event does not match its authorized subscription." Note the word authorized. The subscription is not just routing metadata; it is the authorization record for the event. If an operation is renamed or re-versioned, old events do not silently follow it — they fail the binding check. If a subscription is paused, in error, or expired, its events stop being accepted regardless of what the provider delivers.

Deduplication before enqueue

Providers redeliver. Webhooks retry after timeouts, pollers re-fetch overlapping windows, and schedulers refire after restarts. The trigger service handles this with content-addressed deduplication keyed on exactly four fields: tenant ID, subscription ID, trigger kind, and provider delivery ID. These are joined and hashed with SHA-256, and the resulting key decides whether an event is new.

The ordering inside accept is the part worth studying. When an event arrives and passes the binding check, the service computes the dedupe key and looks it up. If the key is already recorded, it returns the existing event with duplicate: true and enqueues nothing. If the key is new, it builds a CanonicalTriggerEvent — a evt_ identifier, the tenant and subscription, the source kind, the provider delivery ID, an idempotency key of the form trigger:<hash>, the occurrence timestamp, and the opaque payload — then records the event before calling the enqueuer. The inline comment explains why: recording first means retries after an enqueue acknowledgement cannot duplicate a run.

This record-then-enqueue order is the correct shape for at-least-once delivery. If the process crashes between recording and enqueueing, the event is recorded but no run was requested — a gap the reconciliation path (discussed below) exists to notice. If the enqueue succeeds and the provider redelivers, the dedupe key catches the replay. What the order cannot survive, in the inspected implementation, is a process restart: the dedupe map is an in-memory Map, so recorded keys vanish with the process. The design anticipates this — the class comment says production must back these records with a database migration — but the inspected code does not implement it. Any claim about durable exactly-once behavior would need that migration and its tests, neither of which is in scope here.

The idempotency key deserves a final note. It is derived from the same hash as the dedupe key and travels with the enqueue request, so the downstream run layer can apply its own dedupe even if the trigger layer's record is lost. Keying on the provider's delivery ID means the guarantee is only as good as that ID's stability: if a provider changes delivery IDs across retries, the trigger layer sees distinct events. The inspected code does not — and cannot — fix an unstable provider identifier.

Three trigger kinds, three arrival paths

The TriggerKind type allows webhook, poll, and schedule. Webhooks arrive through accept directly: the caller presents the event's claimed binding plus the delivery ID, timestamp, and payload, and the service validates, dedupes, and enqueues. Polling has its own dedicated path with stronger ordering guarantees.

The poll method takes a subscription ID, the caller's current cursor, a batch of events, and the next cursor. It enforces two preconditions: the subscription must be a poll subscription, and its stored cursor must exactly equal the presented cursor. Otherwise it throws "Polling cursor is stale or unauthorized." Each event in the batch then goes through the standard accept flow — full binding validation and dedupe — and only after every event is accepted does the stored cursor advance to the next cursor.

Cursor-advance-last is the poll-path equivalent of record-before-enqueue. If processing fails partway through a batch, the cursor does not move, so the next poll re-presents the same window; already-accepted events are absorbed by dedupe, and unprocessed ones get their turn. The strict cursor equality check also serializes pollers: two concurrent poll attempts with the same cursor cannot both advance, because the second finds a moved cursor. What the inspected code does not show is who calls poll, how often, or how delivery IDs are constructed from polled items — the provider-side fetching and the scheduling of poll ticks are outside this file.

Schedules are the thinnest path in the inspected code. A subscription can carry a schedule with nextRunAt, a timezone, and a misfire policy of skip or fire_once. The service exposes dueSchedules, which lists active schedule-kind subscriptions whose next run time has passed. But listing is all it does: there is no timer, no firing logic, no misfire handling, and no update of nextRunAt after a due schedule is observed. The MisfirePolicy type and the overlapMs field are stored on the record and never consumed in this file. A reader should treat schedule support here as data modeling plus a due-query — the policy and mechanics of actually firing scheduled runs live elsewhere or are not yet implemented in the inspected scope.

Lifecycle: provision, renew, remove, reconcile

Provider-side subscriptions expire, so the service manages their lifecycle through the injected provider interface. provision was covered above. renewDue scans all subscriptions for active ones whose renewal timestamp has passed, calls the provider's renew, and records the new timestamp. Renewal failures are contained per subscription: a failed renew marks that subscription error and moves on, returning only the IDs that renewed. Marking error also stops event acceptance for that subscription, since accept requires active status — a failure in renewal propagates to a halt in triggering, which is the fail-closed direction.

remove deletes the provider-side subscription when one exists, then marks the local record paused rather than deleting it. Pausing preserves the binding, cursor, and event history for inspection or reactivation while guaranteeing no further events are accepted. This is consistent with the file's general posture: state transitions restrict triggering rather than destroying evidence.

Two query methods support external supervisors. reconcile returns active subscriptions that need attention: those with an overdue renewal or an overdue scheduled run time. dueSchedules narrows that to schedule-kind subscriptions at or past their next run time. Both are pure listings — they change nothing. Something else must call them on a tick, decide what to do, and perform the renewals or enqueues. The inspected file draws the boundary there: it can tell a supervisor what is due, but the supervisor and its loop are not in this scope.

A worked example

To make the flow concrete, here is an illustrative walkthrough of the inspected code path — a synthetic example traced through the real method logic, not a production observation. Suppose a tenant connects a workflow to a provider webhook. Setup creates a subscription binding tenant t_7 to workflow w_12 version 3, operation op_sync version 1.4.0, connection c_9, kind webhook. Provisioning registers with the provider and stores the provider subscription ID and renewal time.

When the provider delivers event dlv_1001, the caller invokes accept with the full claimed binding. The service checks each field against the stored subscription, hashes t_7, the subscription ID, webhook, and dlv_1001 into a dedupe key, finds no existing record, stores a canonical event with idempotency key trigger:<hash>, and enqueues a run request naming the workflow, version, event, and key. The provider's retry of dlv_1001 recomputes the same key, hits the stored record, and returns duplicate: true with no second enqueue. If the tenant pauses the subscription and the provider delivers dlv_1002, accept throws the binding-mismatch error before any dedupe or enqueue. If operation op_sync ships version 1.5.0 and the subscription still names 1.4.0, events claiming the new version fail the check — version drift halts triggering instead of silently rerouting it.

None of this involves executing workflow logic. The run layer receives a validated, deduplicated request and decides what happens next. That is the separation the file's header promises, visible end to end.

What the inspected code does not prove

Caveats first, because they constrain every claim above. The inspected TriggerService stores subscriptions and events in process-local Maps with an in-memory sequence counter. A restart loses subscriptions, dedupe records, cursors, and renewal state. The class comment explicitly limits this adapter to tests and local development and points to a database migration for production — a migration I did not inspect and therefore cannot describe. Do not read this article as evidence of production-durable scheduling, persistent dedupe, or continuous long-running missions. It is evidence of a boundary design and its local behavior.

The certification context reinforces the limit. The accompanying workflow product report, P07-FINAL-CERTIFICATION-REPORT.md, returns a verdict of NOT_CERTIFIED: the repository state cannot support production certification, with an incomplete checkpoint sequence, unavailable connection-matrix evidence, and unavailable provider canary authority. That report's scope is the production certification question, not the trigger code's local correctness — but it explicitly blocks any claim that the inspected automation surface is production-ready. I treat the two sources together: the code shows how the boundary is designed to behave, and the report forbids presenting that design as a certified production system.

There are also narrower gaps inside the file itself. Webhook signature verification is absent — accept validates the claimed binding against the stored subscription but performs no cryptographic check that the provider sent the payload. Poll fetching, schedule firing, and the supervisor loop that would call reconcile and dueSchedules are all outside the file. The overlapMs and MisfirePolicy fields are stored but unconsumed. The payload is an opaque record with no schema validation. Each gap is a reasonable place for adjacent layers to contribute, but none of those layers was in my assigned sources, so I make no claims about them.

Why the separation matters

Even with those limits, the design has a clear rationale. Binding validation before dedupe means unauthorized or misrouted events cost nothing and leave no trace in the run layer. Dedupe before enqueue means provider retries are absorbed at the cheapest point. Record-before-enqueue and cursor-advance-last orderings push the remaining failure windows toward gaps a supervisor can detect rather than duplicates it cannot undo. And narrow injected interfaces — one enqueue method, three provider lifecycle methods — keep the trigger layer testable without providers or runners: the process-local adapter exists precisely so tests can exercise the boundary logic directly.

The status model supports the same fail-closed posture. Renewal failure marks error, removal marks paused, and both halt acceptance. Stale poll cursors throw instead of merging. Version drift throws instead of rerouting. At every decision point the inspected code prefers stopping the trigger over guessing what the event meant.

That posture is what makes the trigger boundary worth separating from execution in the first place. Execution layers deal in retries, partial progress, and recovery — expensive machinery that should only engage for events known to be authorized and new. The trigger layer's job is to establish exactly that, cheaply, before anything heavier begins. The inspected code shows one concrete way to draw the line: validate the binding, hash the delivery, record first, enqueue narrowly, and leave running the workflow to someone else.