How Ethen Builds Public Model Pages
A public model page is the last step in a longer chain: a route entry, a stored record, an acceptance decision, and a verified read.
A public model page is the last step in a longer chain: a route entry, a stored record, an acceptance decision, and a verified read.
This article traces the Ethen model library architecture from URL pattern to rendered view, using only the current web route table, the model-page server module, and the model-intelligence authority map. It is an implementation explainer. An INDEXABLE decision or an accepted manifest entry does not by itself prove that a page is publicly deployed, reachable, or kept current. This article covers only the web model-page path in the inspected files; it does not describe Chat, Studio, Research, Designer, Founder, Computer, Desktop, or Code behavior.
The route is a promise, not a page
The web route table defines the address shapes a model page can occupy. The relevant entries are models for the library index, models/image-video for a media-specific listing, and models/:publisher/:slug for an individual model detail page. Parallel discovery surfaces are defined under model-intelligence, including model-intelligence/models, model-intelligence/providers, model-intelligence/benchmarks, model-intelligence/leaderboards, and model-intelligence/categories, each with a collection route and a :slug detail route.
Three adjacent route groups matter for understanding delivery. The model-sitemaps/site.xml and model-sitemaps/:file routes define sharded sitemap output. The api/mi/v1/models, api/mi/v1/models/:slug, and api/mi/v1/health routes define versioned model-intelligence API routes and a health route. The docs subtree, with its index, search index, raw, and wildcard entries, is a separate verified documentation destination. A final wildcard route maps to a catch-all handler.
None of these route declarations contains model data. They say where a request goes. Whether a specific publisher-and-model address returns a useful page depends on a generated manifest, a backing object, and a publication decision described below.
Identity comes first
Before any prose or metadata reaches the page, the pipeline checks identity. The server module loads a generated JSON manifest with schemaVersion: 2, a sourceManifest string, a count, and a routes map. At module load it rejects the whole manifest unless the schema version is exactly 2 and the declared count equals the number of entries. That strict startup check means a truncated or mismatched manifest fails loudly rather than serving a partial library.
Each manifest entry carries the fields the rest of the pipeline relies on: id, key, sha256, accepted, publisher, name, publisherSlug, modelSlug, route, schemaVersion, plus optional catalog and media hints, a publicationState, and an updatedAt timestamp. Startup validation also requires every entry to be accepted, requires its route to equal /models/ plus its manifest key, requires the canonical-path helper to agree with that route, requires the checksum to be 64 lowercase hex characters, and requires both the record id and the route to be unique. A duplicate id, a duplicate path, or a malformed hash invalidates the manifest.
URL matching itself is normalized. The normalizeRouteSlug helper decomposes Unicode, strips combining marks, trims whitespace, lowercases the result, replaces every run of non-alphanumeric characters with a single dash, and removes leading and trailing dashes. The resolver applies that normalization to both the publisher segment and the model segment, then looks up the combined key in the manifest map. Display names can therefore contain capitals, spaces, or punctuation while the lookup key stays in a deterministic dashed form.
The stored record must then agree with the manifest on three points. The record's identity.repo_id must equal the manifest entry's id. The record's source.repo_id must also equal that same id. And the record's schema_version must equal the manifest entry's schemaVersion. The route must additionally match /models/ plus the entry's publisher and model slugs. If any of those comparisons fails, the record cannot be treated as the page the manifest entry names, no matter how complete its editorial text may be.
Acceptance is one decision point
The module names resolveModelPublicationState the one publication decision point, and its comment states the governing rule: manifest authority is necessary but not sufficient, and editorial and SEO text are never inputs. The function takes the fetched record, the manifest entry, and a backing state of VERIFIED, MISSING, INVALID, or TRANSIENT.
The early exits handle authority and storage health. An entry that is missing or not accepted returns NOINDEX_NOT_ACCEPTED. A transient backing failure returns TRANSIENT_BACKING_FAILURE. A missing or invalid backing object returns BROKEN_BACKING_OBJECT. A malformed manifest checksum returns NOINDEX_UNVALIDATED. Only after those gates does the function compare identity, source, schema, and route as described above.
The publication-state vocabulary is explicit. Besides INDEXABLE, the module defines NOINDEX_UNVALIDATED, NOINDEX_NOT_ACCEPTED, NOINDEX_SCHEMA_INVALID, NOINDEX_IDENTITY_INVALID, NOINDEX_ROUTE_INVALID, NOINDEX_NOT_PUBLICATION_CANDIDATE, NOINDEX_THIN_SOURCE, NOINDEX_UTILITY, NOINDEX_DUPLICATE_VARIANT, NOINDEX_ENDPOINT_VARIANT, NOINDEX_PENDING, BROKEN_BACKING_OBJECT, HASH_MISMATCH, and TRANSIENT_BACKING_FAILURE. Five of those are terminal catalog decisions held in a dedicated set: thin source, utility record, duplicate variant, endpoint variant, and pending. When the manifest entry carries one of those five states, the resolver returns that same state only if identity, source, schema, and route all agree; any mismatch falls back to NOINDEX_UNVALIDATED. A non-indexable manifest state outside that set likewise resolves to NOINDEX_UNVALIDATED rather than passing through.
The INDEXABLE path adds a candidate check. The record's publication_candidate block must have index_candidate set to boolean true, a recommended_state of either INDEXABLE or INDEXABLE_AFTER_VALIDATION, and a required_before_indexing array in which every entry mentions a provider term and a freshness term. Concretely, each requirement must contain the word provider and one of refresh, freshness, availability, current, or live. The page is indexable only when identity, source, schema, route, and that candidate permission all hold. Anything else in the INDEXABLE branch resolves to NOINDEX_UNVALIDATED.
Two record shapes, one page view
The pipeline accepts two stored schemas and normalizes both into a single ModelPageView. A record whose schema_version is ethen-media-model-v1 follows the media path; anything else must declare ethen-hf-extract-v1 exactly or fail validation. Both paths re-check that identity.repo_id and source.repo_id equal the manifest id, and both require publication_candidate.index_candidate to be a boolean.
The shared view carries the fields a detail page needs: id, name, publisher, both slugs, canonicalPath, a catalogFamily of hf or media, title, a description capped at 240 characters, an indexable boolean derived from whether the publication state equals INDEXABLE, the full publicationState, tasks, architectures, license, languages, tags, use cases, capabilities, limitations, providers, provenance, evidence, and related-model slots. Identity-derived slugs come from splitting the entry id at its slash and normalizing each half, so the view's address segments always derive from the authority id rather than from free-form display text.
The HF-shaped path extracts hub and model-card detail: pipeline tag as task, architectures, model type, a non-negative finite safetensors.total as parameter count, license from card data, library name, base models from either a single string or an array, languages from card data plus hub tags typed as language, up to 24 hub tags, inference providers with task and status, a provider summary, and a model-card title with an excerpt capped at 640 characters. Its creator is the publisher, and its variants and faq lists are empty.
The media-shaped path is structured around endpoints rather than weights. It reads tasks from the media block, builds use cases from editorial use cases plus best_for strings, collects member endpoints into variants with endpoint id, task, and disposition, reads a delivery string, and exposes SEO frequently asked questions that have both a question and an answer. Its creator is null, its architectures, license, library, and parameter fields are empty, and its tags combine tasks with SEO secondary keywords, capped at 24 entries. When a delivery string exists, the view surfaces it as a single provider row.
One media-only transform deserves attention because it is explicitly presentation-only. The scrubProviderBrand helper removes a provider name from the rendered title and description, turning title suffixes and endpoint-count phrases into neutral wording. The comment above that helper states that stored records, routing metadata, and identifiers are unchanged. Branding cleanup affects what the reader sees, not what the pipeline trusts.
Evidence travels with the text
Every editorial claim on the page can carry its own support. The module defines EvidenceText as a text string plus an evidence string array, and FeatureText adds a label with a Capability default. Short descriptions, overviews, use cases, key features or capabilities, and limitations all use those shapes. The view then unions the evidence references from the summary, the overview, each use case, each capability, each limitation, and the SEO evidence list into one deduplicated, sorted evidence array. Media records additionally fold in evidence attached to FAQ entries.
Provenance is a separate, required block. Every view records a sourceType and sourceUrl as mandatory strings, optional capture and processing timestamps, an optional parser version, the manifest entry's storage key as objectKey, and the manifest's sourceManifest identifier. That combination answers where the underlying record came from, when it was captured and processed, which parser build handled it, and which exact stored object and manifest generation back the rendered page.
SEO fields have bounded fallbacks. The page title prefers the stored SEO title and otherwise combines the model name with the publisher or the generic models suffix. The description prefers the stored SEO description, then the summary or overview text, then a generated sentence naming the model and publisher for HF-shaped records or describing grounded capabilities and endpoints for media records. Descriptions are sliced to 240 characters and model-card excerpts to 640, so overlong source text cannot inflate the rendered head tags or card preview.
Delivery verifies before it renders
Loading a page means reading a backing object from a storage bucket, and the loader treats every stage as fallible. It requires the bucket binding, keeps an in-memory cache of at most 128 entries with a five-minute TTL, and races each bucket read against a four-second default timeout. A missing binding, a timeout, or a storage error surfaces as a 503-class failure; only a proven-missing object surfaces as a 404. The failure taxonomy names each outcome: unknown route, missing binding, proven-missing object, timeout, unavailable storage, invalid JSON, invalid schema, and integrity mismatch.
The integrity sequence is the core of this part of the Ethen model library architecture. After fetching the object body as bytes, the loader hashes those exact compressed bytes with SHA-256 and compares the hex digest against the manifest entry's sha256. A mismatch logs an error and raises an integrity failure without attempting to parse the payload. The loader then detects gzip content by its two magic bytes and decompresses through a standard stream when present, parses the result as JSON, and normalizes the record into the page view. Invalid JSON and schema violations each have their own failure kinds and log entries that include the record id and storage key.
Caching and HTTP headers are fixed values exported by the module. Successful views are cached in memory with first-key eviction once the 128-entry cap is reached, and a dedicated reset function clears the cache. The module exports a cache-control value of public, max-age=60, s-maxage=3600, stale-while-revalidate=86400 alongside a retry-after interval of 30 seconds, pairing a 60-second browser lifetime with a 3600-second shared-cache lifetime and an 86400-second stale-while-revalidate window.
Who owns which fact
The authority map in the model-intelligence module decides which system may originate each fact the library displays. Under contract version mi-authority-r1, model-intelligence owns identity, provider mapping, aliases, family and version lineage, release and deprecation status, capabilities, context and output limits, pricing metadata, benchmark metadata, provenance and freshness, and provider policy. Runtime availability, provider certification, and latency and reliability measurements belong to the gateway runtime. Routing decisions belong to cortex policy. Display formatting belongs to the model-library projection.
That division explains why the page pipeline never invents availability or routing. The page can show metadata, provenance, and formatted text from its own record, but live reachability and selection policy come from other owners. The authority module also pins the only two explicit external provider aliases: one external moonshot identifier resolves to kimi, and one zai identifier resolves to z-ai, with every other provider id passing through unchanged. The comment above that map states the rule directly: never infer provider equivalence beyond the listed aliases. A helper exposes the owner for any fact key so consumers can check authority programmatically rather than guessing.
What this pipeline does not prove
The limitations are as specific as the mechanics. An accepted manifest entry is the start of publication logic, not a deployment certificate; the resolver still requires identity, source, schema, route, and candidate agreement before returning INDEXABLE. An INDEXABLE state is a decision about eligibility for indexing, not evidence that a public host serves the page, that a sitemap includes it, or that provider data behind it is fresh. The candidate rule itself demands provider-freshness follow-ups, which means even an indexable record acknowledges that provider state needs rechecking.
Several noindex states are deliberate catalog judgments rather than errors. Thin sources, utility records, duplicate variants, endpoint variants, and pending records are held out of indexing by design when their identity checks pass. Broken backing objects, hash mismatches, and transient failures are distinct operational states with their own handling, not silent fallbacks to stale text. And because runtime availability belongs to the gateway runtime, nothing in the page loader or the manifest proves that any named provider can serve any named model right now.
Readers who want to go further can start from the verified destinations: the library index, the model-intelligence comparison surfaces, and the documentation tree. Those paths, the route patterns, and the authority owners above are the grounded core. Everything else about availability, ranking, or release timing needs its own primary evidence.
The pipeline can therefore be summarized as a chain of agreements: the manifest agrees with itself at startup, the record agrees with the manifest on identity and schema, the candidate block agrees to indexing with provider-freshness conditions, the stored bytes agree with the recorded hash, and each displayed fact stays with its owning system. When all of those hold, the page earns its INDEXABLE state. When any one fails, the module says exactly which kind of failure it was.