From Provider Endpoints to Ethen Model Families
One provider export produced three very different counts. Each count answers a different question — about delivery, identity, and publication.
One provider export produced three very different counts. Each count answers a different question — about delivery, identity, and publication.
A snapshot of Ethen's media-model catalog reports 1,499 endpoints, 491 families, and 128 indexable candidates. None of these numbers contradicts the others. They measure different things: how many delivery endpoints a provider exposes, how many distinct model families those endpoints collapse into, and how many families currently qualify for a public, search-indexable page. This article explains the grouping in between — AI model catalog normalization — and why the three counts must differ.
Three numbers, three questions
Start with what the reconciliation snapshot actually records. The catalog-reconciliation artifact counts 1,499 source endpoints, all accounted for and none unexplained. Those endpoints resolve into 491 canonical families. Of those families, 128 are indexable and 363 are not — 361 held back as thin source material and 2 as duplicate variants. All 491 families are marked enabled for both the Studio and catalog surfaces.
Each figure answers a separate question:
- Endpoints (1,499): how many callable delivery handles exist in the provider export? An endpoint is a route you invoke — one model ID plus one task shape.
- Families (491): how many distinct underlying models do those endpoints represent? A family groups every endpoint variant that belongs to the same model release.
- Indexable candidates (128): how many families have enough substance to justify a public detail page? Publication is a quality gate, not a headcount.
Confusing any two of these produces the classic catalog error: quoting the endpoint count as if it were a model count, or treating the page count as if it were the catalog size. The normalization pipeline exists to keep the three answers separate and reconcilable.
Why one model becomes many endpoints
Providers commonly expose several endpoints for the same underlying model family. The media-models blueprint, which defines the intended catalog design, gives a concrete example: a model such as Nano Banana 2 appears once for text-to-image generation and again for image editing. A video family such as Veo can splinter further — text-to-video, image-to-video, fast and standard tiers, and other endpoint variants. One model release, many callable shapes.
The canonical relationship the blueprint prescribes is one canonical model family to N provider endpoints. Grouping happens deterministically, before any editorial work: parse the export, normalize identifiers and task names, identify the likely developer, and assign each endpoint to its family. The blueprint's worked example resolves two provider-style identifiers into one canonical family with two member endpoints carrying different tasks.
This is also why treating every endpoint as an independent page would be a mistake. Near-identical pages for endpoint variants waste editorial effort, duplicate prose across the site, and create search cannibalization — multiple Ethen pages competing for the same query with nearly the same content. The blueprint is explicit: marketing publishes the canonical family, while Studio retains every useful endpoint. Page count follows identity; execution inventory follows utility.
Developer versus delivery provider
A second normalization step separates two roles that provider exports tend to blur: who built the model and who serves it. In the blueprint's design, fal.ai is the delivery and API provider, but it is frequently not the company that developed the model. The worked example names Google as the developer of Nano Banana 2, with fal.ai as the delivery path. Other developer names the blueprint anticipates include Black Forest Labs, ByteDance, MiniMax, Luma, Runway, Ideogram, Stability AI, and Recraft, among others represented through the same provider.
The canonical record therefore stores both identities side by side — a developer object and a delivery object — rather than deriving authorship from the provider namespace. This matters for every downstream consumer: model pages attribute the right creator, provider hubs aggregate correctly, and Studio can display delivery availability without misattributing the model itself. Getting this wrong would mean crediting the courier for writing the letter.
The pipeline: from export to canonical record
The blueprint lays out a target pipeline that reuses Ethen's existing model-publication architecture instead of building a second, Studio-only system:
provider export
→ deterministic normalization
→ endpoint records
→ canonical family grouping
→ compact factual dossiers
→ editorial enrichment
→ schema validation
→ canonical JSON (gzipped)
→ R2 publication
→ publication manifest
→ Model Library + StudioSeveral properties of this design are worth noting. First, everything deterministic happens in code before any language-model enrichment: parsing, whitespace and URL normalization, pricing extraction, task classification, stable IDs, hashes, and family grouping. The blueprint forbids sending raw scrape content — pricing text, API tables, code samples, boilerplate — for full rewriting. Editorial input is confined to a delta: short descriptions, overviews, capabilities, best uses, and SEO fields.
Second, the pipeline is resumable and versioned. Each record carries its source hash, editorial version, prompt version, and generation timestamp, so a changed snapshot regenerates only affected families. A canary of roughly ten diverse models must pass review before batch enrichment expands.
Third, the publication manifest — not the UI, not page copy, not Studio configuration — is the serving authority. Applications consume canonical structured records; nothing downstream invents its own inventory.
These pipeline stages are design targets described in a September 2026 planning document, not shipped behavior proven by this article's sources. The reconciliation snapshot shows the grouping arithmetic of one processed export; the page-serving code shows how publication authority is enforced at request time. What follows describes each on its own terms.
Why only 128 families are indexable
If 491 families exist, why do only 128 qualify for public pages? Because indexability is a publication decision with explicit states, and most families in this snapshot do not clear it. The reconciliation counts name the gap directly: 361 families carry thin source material and 2 are duplicate variants.
The model-page serving code shows how strictly that decision is enforced. Its publication-state resolver is documented in the source as "the one publication decision point," and manifest authority is necessary but not sufficient: the fetched record must still agree with deterministic identity, schema, and candidate rules before a page renders as indexable. The code defines a full set of terminal noindex states — thin source, utility, duplicate variant, endpoint variant, pending — and passes a record through as noindex only when its identity and schema agree with the manifest authority. Anything that fails those checks lands in an unvalidated state rather than being published on trust.
The blueprint's publication model matches this shape from the planning side: families score as indexable or fall into buckets such as thin, duplicate, endpoint variant, or needs review, and only records in an indexable state enter the main sitemap. Endpoint utility pages, aliases, duplicates, and weak records never become search-indexable pages automatically. The result is deliberate: 128 families strong enough to stand as public pages, 363 held back until — or unless — their source material improves.
One canonical URL per model completes the picture. The blueprint recommends the existing pattern of one detail route per publisher and slug, with category and provider pages linking inward rather than each minting a competing canonical page. The route registry confirms the mechanism exists: a models index, a dedicated image-video landing route, and a dynamic publisher-and-slug detail route served through the manifest-checked page loader.
One catalog, two consumers
The same canonical dataset is designed to serve two audiences with opposite granularity needs. The Model Library publishes families: one page per model, with tasks, modalities, endpoint variants listed as variants, pricing, specifications, and editorial content. Studio consumes endpoints: the model picker, task filters, pricing display, and endpoint routing all need the full 1,499-handle inventory, because a creator choosing between text-to-image and image-editing variants of one family is choosing between endpoints.
Studio is a separate target application from Chat, which stays deliberately limited in scope — but target separation is an architecture statement, not evidence that any particular catalog integration has shipped. The blueprint describes Studio deriving its picker fields, default endpoints, and capability flags from deterministic endpoint metadata wherever possible, rather than maintaining hardcoded model metadata of its own. Whether a given Studio surface already reads from the canonical catalog is a deployment question this article's sources do not answer, and this article makes no claim about it.
What the sources do establish is the authority hierarchy the blueprint prescribes: source snapshot, normalized endpoint record, canonical family record, editorial delta, merged record, R2 object, publication manifest, applications. The serving code corroborates the enforcement end of that hierarchy — manifest-checked routes, checksum-verified objects, schema-validated records — for the model-page consumer. One catalog, one authority, multiple consumers: that is the design principle, with the page-serving half of it visible in implementation.
What enablement does not prove
The most important limitation in the snapshot is also the easiest to miss. All 491 families are marked enabled for Studio and catalog surfaces — but catalog enablement is not execution proof. An enabled flag means the family is admitted to the catalog inventory; it says nothing about whether anyone has successfully invoked its endpoints, measured latency, verified output quality, or confirmed current provider availability. The claim limit on this article exists precisely to hold that line: do not read runnable-model counts out of catalog counts.
Related caveats deserve equal weight:
- The counts are a snapshot. The reconciliation artifact records one processed export. New provider snapshots add, change, or remove endpoints, and the blueprint's update design processes only deltas — so every count in this article is already historical the moment the next export lands.
- Unknowns dominate the tail. Of 491 families, 145 are classified unknown — the single largest bucket, ahead of image (159) and video (110). Audio, speech, music, 3D, language, LoRA training, and utility families make up the remainder. A catalog whose largest category is "unknown" is being honest about its classification limits, and any summary that omits that bucket misrepresents the data.
- Thin source is the norm, not the exception. Roughly 261 of 1,499 source rows carried substantial prose, against about 1,432 with pricing. The export is rich in endpoint metadata and poor in description — which is exactly why only 128 families clear the publication bar and why the blueprint routes editorial effort through compact dossiers rather than full rewrites.
- Plans are not shipments. The blueprint is a dated planning document; it proves the design's scope, not its deployment. The page-serving code proves enforcement mechanics, not catalog completeness. Neither source supports launch or availability claims for any product surface.
Counting what matters
Provider endpoints, model families, and indexable pages form a funnel: 1,499 delivery handles normalize into 491 identified families, of which 128 earn public pages. Each stage discards a different kind of noise — variant duplication at the grouping stage, thin and duplicate material at the publication stage — while preserving the full endpoint inventory for the consumers that need it. The catalog is most useful when all three numbers stay visible and each is read for what it measures: delivery breadth, model identity, and publication quality. Anything else is a count in search of a meaning.