Skip to content

EthenEthenEthen

Choosing Image and Video Models by Workflow

Pick the workflow before the model: a durable checklist for task fit, inputs, execution eligibility, and delivery checks.

Pick the workflow before the model: a durable checklist for task fit, inputs, execution eligibility, and delivery checks.

Most image and video model selection goes wrong before any model is compared. Someone picks a name they recognize, writes a prompt, attaches a reference, and only then discovers the route they wanted does not accept reference images, or requires a duration they did not supply, or was never an executable path at all — just a catalog entry. A workflow-first approach avoids that failure mode. Decide what kind of job you are running, check what that job requires as input, confirm the route is actually eligible to execute, and verify what delivery produces. This guide walks through each step with concrete criteria drawn from working implementations, and it stays useful as model names change because the decision structure does not depend on any single model.

Start with the workflow, not the model name

Image and video generation is not one task. In the implementation this guide draws on, the media pipeline distinguishes four workflows: text-to-image, image-editing, text-to-video, and image-to-video. Three of them — text-to-image, image-editing, and text-to-video — travel through one validated command path; image-to-video executes on a separate pinned handler and is explicitly refused on the shared path with a wrong-handler error. That separation is the single most important selection fact: workflows that sound similar can live on entirely different execution machinery, with different inputs, different validation, and different failure modes.

The practical consequence is a simple rule. Name your workflow in one sentence before you look at any model list: "generate a still from text only," "modify a supplied reference image," "generate a clip from text," or "animate a supplied image." If you cannot state which of these you mean, you are not ready to choose a model — because the very next questions (do I need a reference? a duration? a resolution?) have opposite answers depending on the workflow.

A second reason to lead with workflow is that selection surfaces and execution paths are different things. A model picker organizes entries into sections such as Chat, Image, Video, and Voice, and those sections are generation modes, not mutually exclusive provider categories: one entry may legitimately appear under several sections. The picker described in the inspected source even tracks unique-model counts separately from visible per-section entries for exactly this reason. But presence in a picker section is catalog metadata. The source states explicitly that curated entries carry no runtime readiness claim, that no provider connectivity is implied by presence in the picker, and that no generation API is wired into the picker itself. Several catalog entries say so on their face with the note "generation is not connected." So the fact that a name appears under an Image or Video tab answers "does this catalog know about this model?" — it does not answer "can I run this workflow on this model right now?" Only the execution path answers that.

Task fit: what each workflow is for

With that framing, here is how the four workflows differ and what each one fits.

Text-to-image takes a text prompt and produces still images. It accepts no reference images and no video parameters: no duration, no resolution. An aspect ratio may be supplied within a fixed set of supported values. This is the workflow for fresh visuals from a description — illustrations, backgrounds, concept art, placeholder imagery — where no existing pixels need to survive into the output.

Image-editing takes a prompt plus at least one reference image and produces modified stills. The reference requirement is structural, not optional: the command path rejects an image-editing request with zero references, just as it rejects references on workflows that accept none. This is the workflow for retouching, restyling, object changes, background swaps, and any task where the output must preserve part of an input image. If your task has a "keep this, change that" structure, it belongs here, not in text-to-image.

Text-to-video takes a prompt plus a small set of mandatory video parameters: a duration chosen from a fixed enum, an aspect ratio from a fixed set, and a resolution from a fixed set. In the inspected implementation the duration enum is exactly 5, 10, or 15 seconds, the resolutions are 720p or 1080p, and the aspects are 16:9, 9:16, 1:1, 4:3, or 3:4. This is the workflow for generated clips from a description — establishing shots, loops, motion backgrounds — where the length, frame shape, and resolution are decided up front rather than discovered.

Image-to-video animates or extends a supplied image. It runs on its own pinned handler rather than the shared media command path, which means its input contract, validation, and error behavior are defined separately. Treat it as a distinct integration surface: assumptions carried over from text-to-video (or from image-editing) do not automatically transfer.

A durable task-fit checklist, usable against any provider:

  1. Does the output start from nothing (text only) or from supplied pixels? Text only points to text-to-image or text-to-video; supplied pixels point to image-editing or image-to-video.
  2. Is the output a still or a clip? Stills point to the image workflows; clips point to the video workflows.
  3. If a clip, are length, shape, and resolution fixed choices from a menu or free-form values? Fixed menus mean validating against the enum before submitting, not after failing.
  4. Does the workflow you chose execute on the same path you are submitting to? A request sent to the wrong handler fails no matter how good the prompt is.

Note what this checklist deliberately omits: any ranking of which model is "best." This guide makes no best-model claims. Quality comparisons require measured first-party results on defined tasks, which the inspected sources do not contain. Task fit narrows the field to workflows that can structurally do your job; it never crowns a winner.

Inputs: the validation checklist

Every media command is validated before anything else happens — before health checks, before pricing, before enqueueing. The inspected pipeline follows a strict order: validate, then assess route health, then quote, then admit against an approval ceiling. Inputs that fail validation never reach execution, so the cheapest failures to fix are input failures. Here is the input checklist, with the concrete limits from the inspected implementation as examples of the kinds of constraints to verify on whatever route you use.

Prompt. A non-empty prompt is required, with a maximum length (4,000 characters in the inspected code). Prompts that exceed the limit are rejected rather than truncated. Evergreen advice: keep prompts well under any stated maximum, and never assume silent truncation — verify whether your route truncates, rejects, or counts differently (characters versus tokens).

Negative prompt. Support for negative prompts is route-specific, not universal. In the inspected code, only two routes accept one; every other route rejects the request if a negative prompt is supplied, with an explicit "omit negativePrompt" error. The limit where supported is 1,000 characters. The general lesson: optional creative controls (negative prompts, style strengths, guidance scales) are per-route capabilities. Check support on the exact route before including them, because an unsupported-but-present parameter can fail the whole request rather than being ignored.

Reference images. At most five references are accepted in the inspected pipeline, each validated as a locator that must resolve to a canonical address and identifier. Image-editing requires at least one; the text-only workflows reject any. References keep their submitted order through to provider delivery. When you compare routes, ask four questions: are references accepted at all, how many, in what form (uploaded asset, URL, or both), and does order matter?

Seed. Reproducibility controls are bounded: the inspected seed must be an integer from 0 to 2,147,483,647. Non-integers, negatives, and out-of-range values are rejected. If repeatability matters to your workflow — regression tests on prompts, iterative art direction — confirm that the route accepts a seed and that identical seeds actually reproduce outputs on that provider, rather than assuming determinism.

Video geometry. Text-to-video requires all three of duration, aspect ratio, and resolution, each from its fixed enum. Omitting any one of them, or supplying a value outside the enum (a 7-second duration, a 4K resolution), fails validation. Image workflows accept an aspect ratio from a wider set of eight values but reject duration and resolution outright. The evergreen principle: geometry is part of the request contract, not a post-processing preference. Confirm the enums before you design storyboards or layouts around dimensions a route cannot produce.

Provider parameter mapping. Even after validation, your logical inputs are translated into provider-specific bodies, and the translation differs per route. One image route maps aspect ratios to pixel grids (16:9 becomes 1344×768, 4:3 becomes 1184×880, the default square is 1024×1024); an editing route maps them to named sizes (landscape_16_9, portrait_4_3, square); a video route passes duration, resolution, and aspect through as separate fields. Requested image counts are fixed at one per call in the inspected builders, a safety checker is enabled on every route, and the editing route pins its output format. None of these details are visible in the logical command — they emerge from the route's verified schema snapshot. The lesson for selection: two routes that accept the same logical inputs can still differ in output dimensions, count, format, and safety behavior. Read the route's parameter mapping, not just its input list.

Execution eligibility: catalog presence is not a live route

The most expensive selection mistake is treating catalog presence as execution readiness. The inspected sources draw this line with unusual clarity, and it deserves a section of its own.

A media catalog reconciles provider endpoints into model families: the inspected reconciliation counts 1,499 source endpoints resolved into 491 canonical families, with 159 image families and 110 video families among them. But only 128 families are indexable; 363 are held back, 361 for thin sourcing and 2 as duplicate variants. Endpoint counts, family counts, and executable-route counts are three different numbers, and confusing them produces inflated expectations. A catalog can know about fifteen hundred endpoints while offering a handful of verified execution routes. This is normal and healthy — cataloging is cheap, verified execution is expensive — but only if you read the numbers for what each one measures. (These figures describe the reconciliation artifact's own scope at its snapshot date; they prove that counting discipline, not any current catalog size.)

Schema qualification adds a second layer that still does not prove a live route. A route may have a verified schema snapshot — exact enums, required fields, a parameter builder written against that snapshot — and still not be the thing that executes your request. In the inspected pipeline, execution requires a chain of further gates. First, the workflow must resolve to a verified route at all; unknown workflows are refused. Second, the route's handler kind must match the submission path; a route pinned to a different handler is refused with an explicit error even though its schema is fully known. Third, the route must pass a health gate computed over measured durable outcomes: a route whose measured success rate falls below one half, on medium or high confidence evidence, is refused as unhealthy. Fourth, pricing must resolve — an active, unretired pricing version for the provider, model, and capability — or no quote can be produced. Fifth, the quote must fit within an approved ceiling, otherwise the request stops with an approval requirement rather than executing.

Two subtleties in that chain are worth internalizing because they generalize. The health gate treats "no evidence" differently from "bad evidence": thin or stale samples still route (when the route is the sole verified option) but attach a disclosed low-confidence note to the verdict and receipt. Empty history is reported as unknown confidence, never as healthy. And pricing is versioned and capability-scoped: a price for one capability on one model says nothing about another capability, and a retired pricing version behaves the same as no pricing at all.

A durable eligibility checklist, in gate order:

  1. Workflow resolution. Does this workflow resolve to a verified route on the path you are using? Unknown workflows and wrong-handler submissions fail first.
  2. Schema conformance. Do your inputs satisfy the route's exact snapshot — enums, required fields, reference rules? Anything the snapshot does not support is rejected, never silently dropped.
  3. Health. What is the route's measured success rate, on how many samples, and how fresh? Below a coin flip on solid evidence means do not offer the route; thin evidence means proceed only with the uncertainty disclosed.
  4. Pricing. Is there an active pricing version for this provider, model, and capability? Video quotes may scale with duration (a base credit cost plus a per-second rate times duration), so confirm the pricing shape, not just a number.
  5. Approval. Does the quote fit the approved ceiling? If not, the correct outcome is a held request awaiting approval, not an execution.

If any gate lacks evidence you can inspect, say so explicitly rather than assuming it passes. An unverified gate is a reason to qualify the recommendation, not to block the write-up — but it must be visible.

Delivery checks: receipts, reservations, and recovery

Selection does not end when a request is accepted. The delivery half of the workflow determines whether you can audit what ran, what it cost, and what to do when the outcome is uncertain. The inspected pipeline returns a structured receipt with every admission: provider, model, capability, adapter version, and schema snapshot, plus a routing receipt recording how the route was chosen. Treat equivalent fields as your minimum delivery checklist on any platform: which provider and model actually executed, under which capability and adapter version, against which schema snapshot, and why that route won. If a platform cannot tell you these after the fact, you cannot reproduce the run or explain its cost.

Execution is durable and idempotent. Requests enqueue with a reservation key, and replays of the same key return the existing reservation rather than double-charging: the response carries an explicit replayed flag. Health sampling, in turn, reads only terminal job states — completed, failed, cancelled, and the various uncertain-outcome states — and counts only completed runs as accepted output with settled credits. These two facts combine into practical delivery checks: confirm idempotency behavior before retrying (a safe retry needs a stable key, not a fresh request), and confirm how your platform classifies non-completed outcomes, because retries, refunds, and health accounting all hinge on that classification.

For video especially, verify delivery attributes against what you requested: duration, resolution, aspect, and output location. Geometry enums constrain the request, but the delivered file is what your downstream workflow consumes — a mismatch caught at delivery review is cheap, while one caught in editing is not.

Three worked examples

Example 1: a marketing still from a text brief. The workflow is text-to-image: no reference images exist, and the output is a still. Inputs are a prompt within the length limit, an optional aspect ratio from the supported set, and an optional seed for iteration. Eligibility checks confirm the workflow resolves to a verified route on the submission path, the aspect value is inside the enum, the route's health evidence is acceptable, pricing resolves, and the quote fits the ceiling. Delivery review confirms the receipt fields and the output dimensions implied by the aspect mapping. If the brief later gains a "keep this product shot, change the background" requirement, the workflow changes to image-editing — a different route with a mandatory reference — rather than a better prompt on the same route.

Example 2: restyling an existing illustration. The "keep this, change that" structure makes this image-editing from the start. The reference image is mandatory input, validated as a resolvable locator; up to five references may be supplied in order. A negative prompt is included only if the exact route supports it — otherwise it fails the request. Seed handling supports iterative comparison across attempts. The eligibility chain is otherwise identical: verified route, schema conformance, health, pricing, approval. The delivery receipt's schema snapshot matters here because editing behavior (output format, size mapping) is route-defined.

Example 3: a ten-second clip for a landing page. The workflow is text-to-video, which immediately fixes the required inputs: duration 10 (inside the 5/10/15 enum), an aspect from the video set, and a resolution of 720p or 1080p. All three are mandatory; the request fails without any of them. Pricing may combine a base cost with a per-second rate, so the quote — and therefore the approval ceiling — depends on the duration choice. Delivery review checks the clip's actual duration, resolution, and aspect against the request, plus the standard receipt fields. If the plan changes to animating an existing still, the workflow becomes image-to-video on its separate pinned handler, and the input contract must be re-verified from scratch rather than adapted from the text-to-video enums.

Limitations: what this guide cannot tell you

Honest selection requires stating what the evidence does not support. First, nothing here ranks models. The inspected sources contain catalog structures, validation rules, routing gates, and counting discipline — not measured quality comparisons — so any "best model for X" claim would be unsupported. Quality verdicts need defined tasks, fixed evaluation conditions, and first-party results; until those exist for your task, treat quality claims from any source as unverified.

Second, catalog and schema evidence never proves a live route. Curated picker entries explicitly disclaim runtime readiness; verified schemas still sit behind handler matching, health gates, pricing, and approvals. When you evaluate a platform, ask to see the execution chain — resolved route, health evidence, active pricing, admission decision — not just the catalog page.

Third, the concrete values in this guide (enums, limits, counts) describe the inspected implementations at their snapshot scope. Providers change enums, catalogs reconcile differently over time, and health evidence moves with every run. The checklists are the durable part; the example values illustrate the kinds of constraints to verify, not a standing configuration. Re-check enums and limits on the route you actually use, every time the route or its snapshot version changes.

Fourth, product-surface boundaries are design facts, not availability claims. The picker-versus-execution distinction, the separate media handlers, and the multi-app product map describe how responsibilities divide; they do not prove what has shipped or what you can access today. Selection guidance that respects this boundary stays accurate across releases, while guidance that conflates a target architecture with a live product expires at the next reorganization.

A one-page selection checklist

For quick reference, the full workflow in order:

  • Name the workflow in one sentence: text-to-image, image-editing, text-to-video, or image-to-video. If it has a "keep this" component, it needs references and belongs on an editing or animation path.
  • Validate inputs before anything else: prompt present and within limits; negative prompt only where supported; references present exactly where required and absent where forbidden; seed in range; video geometry complete and inside the enums.
  • Confirm eligibility in gate order: verified route, matching handler, schema conformance, acceptable health evidence, active pricing, quote within the approved ceiling.
  • Verify delivery: receipt fields (provider, model, capability, adapter version, schema snapshot), routing rationale, replayed-versus-new execution, and output attributes matched against the request.
  • Record the limits: note every gate you could not inspect, every enum you did not re-verify, and every quality claim without measured evidence — and treat those notes as part of the decision, not footnotes to it.

Model names will keep churning. Workflows, input contracts, eligibility gates, and delivery receipts change far more slowly, and they are where selection decisions actually succeed or fail. Choose by workflow, verify each gate in order, and let the checklist — not the catalog's longest list — make the call.