Skip to content

EthenEthenEthen

Building Durable Image and Video Jobs in Ethen Studio

How one command pipeline turns a media prompt into a quoted, reserved, and reconcilable job — and where the current evidence stops.

How one command pipeline turns a media prompt into a quoted, reserved, and reconcilable job — and where the current evidence stops.

Generating an image or a video clip looks like a single action from the user's side: write a prompt, get a result. Inside Ethen Studio, that action passes through a pipeline with five distinct stages — validation, health gating, quoting, reservation-backed admission, and delivery with reconciliation. Each stage can refuse the job, and each refusal names its reason. This post walks that pipeline as implemented in Studio's fal media command path, following the September 20 Job 03 audit report and the command code itself. Studio is a separate target application from Chat, Research, Designer, and Founder, and nothing here is a launch or availability claim: it describes an implementation and its test evidence, including the limits the audit states explicitly.

Four qualified workflows, one gate

The Job 03 report qualifies four durable workflows, each pinned to a verified provider endpoint with a schema snapshot and a handler:

  • Text-to-image runs on fal-ai/flux/dev through the fal-media handler.
  • Image-editing runs on alibaba/qwen-image-3/edit through fal-media.
  • Text-to-video runs on wan/v2.6/text-to-video through fal-media.
  • Image-to-video runs on fal-ai/wan-i2v through the pinned fal-video handler.

Two details matter here. First, the gate for execution is Studio qualification — the verified route table in fal-routes.ts — not the source catalog's disposition. A model appearing in a catalog does not make it runnable; only a verified route with an evidence reference does. Second, image-to-video stays on the pinned handler, and the fal-media command refuses it outright with a WRONG_HANDLER error. Handler assignment is explicit, and the wrong door stays shut rather than guessing.

The shared provider seam underneath (providers/fal.ts) exposes parameterized submit, poll, fetch, cancel, and download operations plus an endpoint guard, extended with additive branches only — the report notes pinned behavior stayed byte-identical. That discipline matters because durability work touches code paths that already serve production-like traffic in tests; additive branches keep the new pipeline from changing what already works.

Validation: reject what the snapshot does not know

The command pipeline begins with validateFalMediaCommand, and its governing rule is strict: provider parameter builders follow the verified schema snapshots exactly, and anything the snapshot does not support is rejected, never silently dropped. A parameter the provider would ignore is worse than an error — it is a promise the system cannot keep.

The concrete rules show what this strictness looks like in practice:

  • The prompt is required, trimmed, and capped at 4,000 characters.
  • The negative prompt is capped at 1,000 characters and accepted only on the two routes whose snapshots support it — the Qwen image-edit route and the Wan text-to-video route. Sending one to Flux is a validation error, not a quiet omission.
  • The seed must be an integer from 0 to 2,147,483,647.
  • At most five reference images are accepted. Image-editing requires at least one; the other two fal-media workflows accept none at all.
  • Text-to-video is fully enumerated: duration must be one of 5, 10, or 15 seconds, resolution must be 720p or 1080p, and the aspect ratio must be one of 16:9, 9:16, 1:1, 4:3, or 3:4. Image workflows accept no duration or resolution, and their aspect ratios come from a fixed eight-value grid.

The provider input builders then translate validated commands into exact queue bodies. Flux gets an image_size pixel pair mapped from the aspect ratio (1344×768 for landscape shapes, 768×1344 for portrait, and so on) with num_images fixed at 1. The Qwen edit route receives image_urls as an array, an image_size token such as landscape_16_9 or auto, PNG output, and the negative prompt only when present. The Wan video route receives duration as a string — matching the snapshot's string enum verbatim — plus resolution and aspect ratio. Every body sets enable_safety_checker to true. Because the builders follow snapshots exactly, a snapshot change is a deliberate, reviewable event rather than drift.

The health gate: measured routes or disclosed uncertainty

Before money is discussed, the pipeline checks whether the route has earned traffic. assessFalRouteHealth applies a simple bar: a route whose measured success rate falls below 50% at medium or high confidence is refused with UNHEALTHY_ROUTE. A route failing more often than a coin flip is not offered to users.

The more interesting case is the unmeasured route. When evidence is thin or stale — unknown or low confidence, or samples past their freshness window — the pipeline still routes, because there is exactly one verified route per workflow and no alternative to prefer. But it discloses that: the verdict carries a note naming the route, the sample count, the confidence level, and whether the evidence is stale. The note travels onto the receipt, so downstream records show the job ran on disclosed uncertainty rather than on a claim of health.

Health samples come from durable execution history, not from anecdotes. collectFalRouteSamples reads bounded job history for the project, keeps only terminal statuses — completed, failed, cancelled, dead-letter, timed out, indeterminate, escalated, halt-unsafe — filters to the route's endpoint, and treats completed runs as accepted output with their quoted credits settled. Empty history means unknown confidence, and the code comment says explicitly that callers must never read an absence of failures as health. That distinction — no evidence of failure versus evidence of success — is the kind of thing durability engineering gets right by writing it down.

Quotes and reservations: price before admission

With validation passed and health assessed, the pipeline prices the job. lookupFalPricing reads the active global pricing row for the provider, model, and capability from studio_pricing_versions — organization global, version v1, not retired. Credits must be a finite non-negative number or the lookup returns null, which fails the command closed: without a pricing row, there is no quote, and without a quote, no admission. The audit flags this directly — no pricing rows were seeded for the three new routes, so commands fail with SETUP_REQUIRED until seeding happens. Fail-closed beats inventing a price.

Quoting itself is straightforward. Image jobs cost the row's flat credit amount; text-to-video adds a per-second rate multiplied by duration when the row carries one. admitFalQuote then compares the quote against an approved ceiling and refuses with APPROVAL_REQUIRED when the quote exceeds it. Pricing, quoting, and approval are three separate steps with three separate failure modes, which makes each one auditable.

Admission is where durability begins in earnest. enqueueFalMediaCommand ranks the route (a single-entry ranking that fails internally if the verified route is not the winner), builds the execution receipt, and enqueues through enqueueWithReservation with an idempotency key prefixed fal-. The job payload carries everything the worker and the reconciler will need later: provider and model identifiers, workflow, endpoint, prompt, negative prompt, reference canonicals and reference IDs, seed, duration, aspect ratio, resolution, the actor, the reservation key, quoted credits, the pricing version, the capability version, the health note, and both receipts. Retrying the same idempotency key replays the existing reservation instead of creating a second charge — the response reports whether the reservation was replayed, so callers can distinguish a new job from a recovered one.

Delivery: receipts, markers, and settle-after-asset

The worker side — the fal-media handler — is where the audit's lifecycle language gets concrete. The handler practices reservation admission (no reservation, no execution), receipt agreement (the worker confirms the terms the command promised), and dispatches under a falmedia: marker. Three recovery behaviors stand out:

  • Never-resubmit recovery. If the worker restarts or the outcome is unclear, it does not blindly submit a second provider request. Resubmission is how one prompt becomes two charges.
  • Upstream cancel with re-poll confirmation. Cancellation asks the provider to stop and then re-polls to confirm what actually happened, rather than assuming the cancel landed.
  • Late-success-after-cancel guard. If a result arrives after a cancel was requested, the guard prevents it from being treated as a normal delivery.

Settlement follows the settle-after-durable-asset rule: the job is not marked complete until the output asset is durably stored. A provider response that cannot be turned into a stored, retrievable asset is not a delivery. Job reads stay consolidated on the payload-generic readImageJob, and fal-media payloads carry provider and model identifiers so canonical health sampling sees them — the telemetry loop closes because the identifiers are part of the payload, not reconstructed afterward.

Reference handling deserves its own note because it spans command time and worker time. Tenant asset: canonicals resolve to fresh signed URLs at worker time, so URLs cannot go stale between admission and execution. Remote URLs are validated twice — once at command time, once at worker time — and private, credentialed, and cleartext targets are refused both times. Double validation with SSRF-safe handling means a reference that was safe at admission is re-checked at the moment of fetch, closing the window where a URL's meaning could change.

Uncertain outcomes: reconcile or escalate

The hardest part of durable execution is the job whose outcome is genuinely unknown — the provider never answered, the worker died mid-poll, the response was ambiguous. The pipeline treats "indeterminate" as a first-class status with its own path. Retry re-queues failed jobs on the existing reservation, but indeterminate jobs must pass through reconciliation first; you may not retry what you have not yet understood.

Reconciliation (fal-reconcile.ts) resolves indeterminate jobs against fal queue truth: it asks the provider's own queue what happened and settles the job as completed, failed, or escalated. Escalation is the honest terminal state for jobs that cannot be resolved automatically — flagged for human attention rather than forced into a success or failure they did not earn. A project-scoped sweep endpoint drives this periodically, and it leaves non-fal-media jobs untouched, so the new pipeline cannot disturb jobs owned by other handlers.

This is the reconcile-or-escalate settlement the audit's verdict names. The design accepts that some outcomes stay uncertain after every automated check, and it gives those jobs a visible, bounded destination instead of letting them linger as silent unknowns.

What the evidence does and does not prove

The audit reports a PASS with substantial stated verification: 30 of 30 runtime tests in the new behavioral suite, plus green companion suites (32 catalog-controls tests, 50 image-durable, 70 video-delivery, 42 quota-admission, 4 fal-durable checks), clean TypeScript compilation for both the Studio app and the platform worker, and zero eslint errors or warnings on touched files. That is meaningful evidence about the code paths exercised — validation branches, quote math, reservation replay, handler lifecycle under injected failures.

The claim limit is equally important, and the report states it plainly. No live provider calls were made; provider behavior is covered by seam guards and injected-failure handler tests, not by observed fal responses. Pricing rows for the three new routes were not seeded, so the quote path fails closed until that setup lands. And one procedural note: the Job-01 audit directory vanished mid-Job-2, which is why this report lives in-repo. A dated report proves only its scope — these results describe the tested behavior of this pipeline on September 20, not a guarantee about provider behavior, production traffic, or future routes.

Readers comparing this with sibling coverage should keep the boundaries straight: this article owns media job execution, while catalog normalization and workflow selection belong to their own write-ups. Execution eligibility, model discovery, and task fit are related questions with different evidence, and merging them would blur what each one actually established.

A pipeline that says no

The through-line of this design is refusal with reasons. Unknown workflows, wrong handlers, unsupported parameters, unhealthy routes, missing prices, exceeded ceilings, unsafe references — each gets a named error at the earliest stage that can detect it. Durability here is not just surviving failures; it is structuring the pipeline so that every commitment (a quote, a reservation, a delivery) rests on checks that already passed, and every uncertainty (thin health evidence, an indeterminate outcome) is labeled and routed rather than hidden. That is what makes the jobs durable: not that nothing fails, but that each failure has a stage, a name, and a next step.