What 1,499 AI Endpoints Taught Us About Model Catalog Design
Good AI model catalog design starts with identity, not presentation. When Ethen reconciled a provider export of media models, 1,499 callable endpoints collapsed into 491 distinct model families, and only 128 of those families had enough substance to deserve a public page. Of the 491 families, 145 could not yet be classified by modality — the second-largest bucket after image. Working through those numbers taught us ten lessons: count what you mean; establish identity before editing anything; publish one page per family; treat publication as a quality gate; show unknown as unknown; invest in structured metadata over prose; never confuse catalog presence with a runnable model; do deterministic work before any language-model enrichment; expect snapshots to age; and give different consumers different granularity. This article explains each lesson for anyone building a catalog of AI models.
Good AI model catalog design starts with identity, not presentation. When Ethen reconciled a provider export of media models, 1,499 callable endpoints collapsed into 491 distinct model families, and only 128 of those families had enough substance to deserve a public page. Of the 491 families, 145 could not yet be classified by modality — the second-largest bucket after image. Working through those numbers taught us ten lessons: count what you mean; establish identity before editing anything; publish one page per family; treat publication as a quality gate; show unknown as unknown; invest in structured metadata over prose; never confuse catalog presence with a runnable model; do deterministic work before any language-model enrichment; expect snapshots to age; and give different consumers different granularity. This article explains each lesson for anyone building a catalog of AI models.
Key takeaways
- Endpoints, families and pages are different counts. Mixing them up is the classic catalog error.
- Identity comes first. Group variants into families before writing a word.
- One canonical page per family. Variant pages compete with each other and help no one.
- Publication is a quality gate, not a headcount. Most families did not qualify.
- Unknown is a real category. Hiding it misrepresents the catalog.
- In a catalog is not the same as runnable. Execution needs its own verification.
The numbers behind the lessons
The lessons come from a single reconciliation snapshot of Ethen's media-model catalog, published in September 2026 in From Provider Endpoints to Ethen Model Families. That post explains the mechanism. This one draws out the design lessons. Figure 1 shows the funnel.
1,499 endpoints. An endpoint is a callable handle: one model plus one task shape. The same model offered for text-to-image and for image editing counts as two endpoints.
491 families. A family groups every endpoint that belongs to the same underlying model release.
128 indexable pages. Families with enough source substance to justify a public page. The remaining 363 were held back: 361 because the source material was too thin, and 2 as duplicate variants.
These are snapshot numbers. The next provider export will change them.
Lesson 1: count what you mean
The first lesson sounds trivial and is the most frequently violated. "How many models do you support?" can mean how many callable endpoints, how many distinct models, or how many have pages. The answers here differ by more than a factor of ten.
A catalog that quotes its endpoint count as a model count overstates its breadth. One that quotes its page count as its catalog size understates it. Either way, readers are misled. The fix is to keep the three counts separate in every document, and to say which one you mean every time a number appears. We return to why model counts are a poor measure of usefulness in Why Model Families Matter More Than Huge Model Counts.
Lesson 2: identity before editing
Before writing any description, comparison or marketing copy, each endpoint needs to be assigned to its family. Grouping is deterministic work: parse the provider export, normalize identifiers and task names, identify the developer, and assign each endpoint to a family. Doing it first means every later step — editorial content, pricing display, routing — attaches to the right thing.
Doing it later is expensive. Descriptions written for individual endpoints have to be merged; links have to be redirected; search engines have to be told which page is canonical. Identity is the cheapest layer to get right at the start.
Lesson 3: one canonical page per family
If every endpoint variant had its own public page, a catalog of 491 families would produce more than a thousand pages, most of them near-duplicates. Near-identical pages waste editorial effort, dilute the signals that help search engines understand a site, and make a catalog compete with itself for the same queries. Search engines' own guidance on duplicate content recommends consolidating duplicates under one canonical page.
So the rule is one public page per family, listing its endpoint variants as variants. The product that runs models — in Ethen's case, Studio — keeps every useful endpoint, because a creator choosing between text-to-image and image editing is choosing between endpoints. Page count follows identity; execution inventory follows utility.
Lesson 4: publication is a quality gate
Of 491 families, only 128 qualified for a public, indexable page. That is not a failure of the catalog. It is the gate working.
A public model page should help someone decide whether to use a model. That needs substance: what the model does, what it is good at, what it costs, what is known about it. For 361 families, the source material did not provide enough. Publishing thin pages for them would have added pages without adding value, and thin pages tend to hurt a site's overall quality signals.
The design choice is explicit publication states — indexable, thin, duplicate, endpoint variant, pending — decided at one point and enforced when pages are served. We describe how Ethen's model pages enforce that decision in How Ethen Builds Public Model Pages. Families that are not indexable are not deleted; they stay in the catalog and can qualify later as their source material improves.
Lesson 5: show unknown as unknown
Figure 2 breaks the 491 families down by modality.
145 families could not yet be classified by modality from the available source material. It would have been easy to guess — most are probably image or video — and produce a cleaner chart. We chose not to. A catalog whose second-largest category is "unknown" is telling the truth about the limits of its sources, and any summary that drops that bucket misrepresents the data.
The same principle runs through Ethen's model information generally: a value appears only when its provenance supports it. See Showing Unknowns in Ethen Model Intelligence and Why Ethen Shows What It Knows—and What It Doesn't.
Lesson 6: invest in structured metadata
The provider export was rich in structured data and poor in prose. Roughly 1,432 of the 1,499 source rows carried pricing; only about 261 carried substantial descriptive text. That shape is typical of provider exports, and it should shape catalog design.
Structured fields — task type, modalities, inputs accepted, pricing, provider, release — can be normalized, validated and compared deterministically. Prose cannot. A catalog built on structured metadata can filter, compare and route reliably even when descriptions are missing. Editorial effort is then spent where it adds most: short overviews and best-use notes for families that qualify for public pages, rather than full rewrites of thin source text.
Documentation research makes a related point. Model cards, proposed by Margaret Mitchell and colleagues in 2019, argued for structured reporting of what a model is for, how it was evaluated and where it falls short. A catalog that captures those fields consistently is more useful than one with long but uneven descriptions.
Lesson 7: catalog presence is not runnable
In the snapshot, all 491 families were enabled for the catalog. That says nothing about whether any given endpoint can be invoked successfully right now, at what latency, or with what output quality.
This distinction matters in both directions. A user browsing a catalog should not assume every listed model can be run. And a product that runs models should never treat a catalog entry as permission to call it. In Ethen Studio, the gate for execution is a verified route — a specific endpoint with a checked input schema and a handler — not catalog membership. We describe that in Building Durable Image and Video Jobs in Ethen Studio.
Lesson 8: deterministic first, language models only on deltas
It is tempting to send an entire provider export to a language model and ask it to write tidy catalog entries. That approach is slow, expensive, hard to reproduce, and risks the model inventing details that sound plausible.
The better order is deterministic work first — parsing, normalization, identifiers, task classification, pricing extraction, family grouping — and language-model enrichment only for a narrow editorial delta: short descriptions and overviews for families that qualify. Each record should carry the source version and the version of the editorial process that produced it, so that a changed snapshot regenerates only what changed. Those are the design targets described in our catalog post, labeled there as planning-document targets rather than proven behavior.
Lesson 9: snapshots age
Every number in this article was already historical the moment the next provider export arrived. Providers add, rename and retire endpoints constantly. A catalog design that assumes a static inventory will drift.
Two habits help. Process changes as deltas, comparing each new snapshot to the last and regenerating only affected families. And date every count you publish. A count without a date invites readers to treat a snapshot as permanent.
Lesson 10: different consumers need different granularity
The same catalog serves different readers. A person comparing models wants families: one page per model, with variants listed. A product that executes work needs endpoints: the full set of callable handles, with their task shapes and inputs. A routing system needs eligibility: which endpoints are healthy, affordable and permitted right now.
Designing one canonical dataset with several views — families for people, endpoints for products, eligibility for routing — avoids maintaining three separate catalogs that drift apart. Each fact needs a clear owner, a principle we describe in Who Owns Each Model Fact in Ethen.
Mistakes we avoided, and one we nearly made
Several tempting shortcuts would have produced a bigger-looking catalog faster.
Publishing every endpoint. A page per endpoint would have multiplied the page count roughly threefold, with most pages nearly identical. It would have looked impressive in a sitemap and served readers badly.
Guessing modalities. Assigning the 145 unclassified families to image or video by pattern-matching names would have tidied the chart and introduced errors we could not see.
Rewriting thin sources with a language model. Generating full descriptions for the 361 thin families would have filled pages with fluent text that the sources did not support.
The mistake we nearly made was subtler: treating "enabled in the catalog" as if it meant "available to use". The two flags look similar in data, and a dashboard that shows only one of them invites the wrong conclusion. Keeping catalog enablement and verified execution as separately named, separately displayed facts was one of the most useful decisions in the whole exercise.
The lessons together
Figure 3 groups the ten lessons.
The identity lessons come first because they determine everything else. Quality lessons decide what is shown and how honestly. Truth-over-time lessons keep the catalog accurate as providers change.
What this means for people choosing models
If you are using a model catalog — Ethen's or anyone else's — a few questions help you read it accurately.
Is this count endpoints, models or pages? Ask which one any headline number means.
Is this model runnable here, now? A listing is not a guarantee.
What is unknown about it? A catalog that shows unknowns is easier to trust than one that fills every cell.
When was this information updated? Model facts age quickly.
For choosing among models, see How to Choose an AI Model, and for browsing Ethen's families, the Model Library.
Tradeoffs and limitations
Strict gates leave gaps. Most families have no public page. Users looking for a specific model may not find a page for it yet.
Unknowns look untidy. A large unknown bucket makes the catalog look less complete than a guessed one would. We prefer accuracy.
Snapshot counts date quickly. The numbers here describe one September 2026 export.
Some design targets are not yet proven. Delta processing and versioned editorial records are described in planning as targets; the published post says which parts are shown in implementation.
FAQ
How should an AI model catalog be organized? By model family, with endpoint variants listed under each family, one canonical page per family, structured metadata for comparison, and clear states for which families are public.
What is the difference between a model endpoint and a model? An endpoint is a callable handle for one model and one task shape. A model family groups all the endpoints that belong to the same underlying model release.
Why do model catalogs have so many duplicates? Because providers expose the same model through several endpoints — different tasks, speeds or tiers. Without grouping, each looks like a separate model.
Does a model appearing in a catalog mean I can use it? Not necessarily. Catalog presence and verified, runnable access are different things.
How many models does Ethen's catalog have? In the September 2026 snapshot: 1,499 endpoints grouped into 491 families, of which 128 had public, indexable pages.
Related reading
- From Provider Endpoints to Ethen Model Families
- Why Model Families Matter More Than Huge Model Counts
- How Ethen Builds Public Model Pages
- Who Owns Each Model Fact in Ethen
- Ethen Model Library
References
- Ethen Blog. From Provider Endpoints to Ethen Model Families (September 2026). https://upcube.ai/blog/from-provider-endpoints-to-ethen-model-families
- Google Search Central. How to specify a canonical URL with rel="canonical" and other methods. https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls
- Mitchell, M., Wu, S., Zaldivar, A., et al. (2019). Model Cards for Model Reporting. Proceedings of FAT\ 2019*. https://doi.org/10.1145/3287560.3287596
- Ethen Blog. How Ethen Builds Public Model Pages. https://upcube.ai/blog/how-ethen-builds-public-model-pages