How Ethen Gateway Chooses an Eligible Model
Eligibility is elimination: the Gateway discards every model that cannot serve a request before it ranks anything at all.
Eligibility is elimination: the Gateway discards every model that cannot serve a request before it ranks anything at all.
The Ethen Gateway does not start by asking which model is best. It starts with a pool of candidates and removes the ones that cannot serve the request: wrong capabilities, unhealthy providers, over-budget estimates, failing telemetry. Only the survivors get scored, and only the top scorer gets selected. When nothing survives, the Gateway returns no selection at all and the caller must not place a provider call. That fail-closed behavior is the single most important fact about Ethen AI model routing, and everything else in the pipeline exists to make the elimination legible.
This article traces that pipeline as implemented in the Gateway's smart-routing code, the model-catalog status logic that decides what each model is allowed to be, and — for contrast — the separate deterministic policy that Ethen Chat's Faros selector uses. Those are distinct code paths solving different problems, and confusing them is the easiest way to misunderstand Ethen's routing.
The pipeline in one pass
The smart-routing module documents its own pipeline in a comment block at the top of the router: a request becomes requirements, requirements merge with an effective policy, and the resulting gates filter a candidate pool through capability checks, live telemetry, budget and cost filters, and provider health. The survivors are ranked, pass through a canary gate, and produce a decision. The decision carries the selected candidate (or null), the full ranking, the list of rejected candidates with reasons, a score with per-component contributions, machine-readable reason codes, and a receipt summarizing the outcome.
Two properties matter upfront. First, the pipeline is deterministic for a fixed input plus a fixed telemetry state: telemetry is the moving part, and everything else is a pure function of policy and catalog state. Second, rejection is a first-class output. Every failed candidate is recorded with a reason, and "no eligible candidate" is an explicit decision with its own reason code, not an exception. When routing refuses, the receipt says why.
Requests enter through a smart alias — a named routing policy — plus a set of requirements, a Gateway route profile id such as text-general, and a request id. Optional inputs narrow the decision further: a project id (which enables project-scoped allowlist and key lookups), a request-level provider restriction, and an effective budget ceiling. An unknown alias is a hard error, not a fallback to a default. That strictness is consistent with the rest of the design: routing never guesses.
Where the candidate pool comes from
Before any filtering, the router builds a candidate pool for the request's route id from Model Intelligence data, pricing, and catalog records. The pool is therefore bounded by what the catalog knows: provider and model identities, capability families, and cost inputs. A model the catalog has never heard of cannot be routed to, no matter how the requirements read.
The route profile matters because "runnable" is defined relative to it. The catalog-status logic treats four text route profiles as the default set — text-general, text-quality, text-creative, and text-reasoning — and records, per provider, which models those profiles reference. A model earns the strongest status only if it is referenced by the current route profiles, among other conditions. Routing eligibility and catalog status are thus two views of the same grounding: the pool a request can draw from is shaped by profiles that someone explicitly wrote, not by an open-ended scan of every model on the internet.
One detail prevents silent mismatch: catalog model ids are provider-prefixed while route profiles carry unprefixed ids, so the status builder strips the prefix before comparing. A code comment notes that skipping this once made the runnable status unreachable. Eligibility comparisons must agree on identity format first.
Capability filters are hard gates, not preferences
The first real elimination stage is the capability filter, which takes the merged requirements — the alias's gates combined with the request's requirements — and divides the pool into survivors and rejected candidates. Anything the request requires, such as vision input, tool support, or structured output, acts as a hard gate: a candidate that lacks a required capability is out, with its reason recorded. There is no partial credit at this stage and no "close enough" promotion later. The scoring stage can only order candidates that already satisfy every requirement, which is why the decision's reason codes always include a marker meaning the winner supports the required capabilities, plus specific markers for vision, tools, and structured output when those were required.
Requirements also carry two quieter gates worth understanding. A certification-level requirement of live restricts the pool to live-certified candidates, and the decision records that restriction in its own reason code. A budget ceiling, discussed in detail below, excludes candidates whose estimated request cost exceeds it. Both are filters, not ranking hints: they remove candidates before scoring rather than merely penalizing them.
Two provider-level constraints apply at the same stage. The request may carry an explicit restriction limiting routing to named providers. Separately, a project allowlist check runs per provider whenever a project id is present — fail-closed, queried per request with no caching, so policy changes take effect immediately. Providers with open circuits also count as unavailable here, so a tripped provider cannot pass on capabilities alone.
The status vocabulary: what a model is allowed to be
Underneath the filter sits a status system that decides what each catalog model is allowed to be in the current environment. The user-facing vocabulary has five states: live, mock, setup-required, unavailable, and unknown. The internal logic maps richer catalog states onto these five, and the mapping encodes several deliberate product decisions.
A model is live only when its internal state is runnable, and runnable is the hardest status to earn. It requires all of the following at once: the provider is configured, the provider adapter is implemented, the exact provider and model pair is referenced by the current Gateway route profiles, and the provider's metadata marks it launch-ready with dated certification evidence. Drop any one condition and the model falls to a weaker state. Models on configured providers that merely lack certification evidence or route-profile membership land in provider-configured, which surfaces as unknown — present, but not something the Gateway claims is executable.
The remaining mappings handle the ways a model can be visible but unusable. A provider whose server-side key is missing in the current environment yields missing-key, surfaced as setup-required. Catalog entries whose capability family the Gateway runtime does not execute — image, video, embedding, rerank, realtime, speech, and transcription — are marked as unsupported modality and surfaced as unavailable, with an explicit note that the runtime currently exposes only read-only catalog discovery for those modalities. Everything else, including entries the repository cannot substantiate as executable, is catalog-only and surfaces as unknown.
Two details deserve emphasis. First, the mock state is narrowly fenced: a provider using a mock fallback surfaces as mock only when the internal state is not missing-key, because the Gateway must never silently mock a production provider just because credentials are absent. Missing keys stay visibly missing. Second, text-compatible families — text, code, reasoning, and long-context — are the only ones that can progress toward runnable; anything else is structurally ineligible for execution regardless of provider health. Readers comparing this with Chat's model picker should note the scope difference: the picker can display models across many modalities, while Gateway execution eligibility is confined to what the runtime actually wires up.
Provider health: keys, adapters, and circuits
Provider health is resolved per candidate after the pool is built and before capability filtering. For each candidate, the router consults three signals: whether a usable key exists, whether the provider's adapter is implemented, and whether the provider's circuit is open. The combination produces an effective health of healthy, degraded, or unavailable, and each value has a different consequence downstream.
Key availability starts from the candidate's existing health marker, but a project id unlocks a second chance: a project-scoped stored key (bring-your-own-key) can make a provider available even without a server-level environment key. That lookup is best-effort — failures resolve to unavailable rather than throwing — and it explicitly excludes the openai-compatible provider, which cannot draw on project keys in this path. Adapter implementation comes from the provider health summary: a provider without a wired adapter cannot be healthy no matter how good its keys look. The circuit breaker provides the third signal, and an open circuit forces unavailable on its own.
The resulting health values behave differently in the pipeline, and the distinction matters. unavailable candidates are excluded by the capability filter — health is a hard gate at that stage. degraded candidates, meaning providers whose adapters are not implemented but which are otherwise reachable, stay in the pool and proceed to scoring, where the health component scores them down. Degraded is therefore a penalty, not a refusal: the router prefers to route around a degraded provider when a healthy alternative exists, but does not pretend the degraded path is equivalent.
The production provider set the allowlist path reasons about is small and explicit: OpenAI, Anthropic, DeepSeek, and the OpenAI-compatible provider. Project allowlist checks test membership against that set for the requesting project. Anything outside it is outside the routing conversation for project-scoped requests, regardless of what the broader catalog lists.
Budgets: the ceiling is a minimum, not a suggestion
Budget handling in the router is deliberately simple. A request may arrive with an effective budget ceiling in dollars — the remaining budget for the current window — and the requirements may carry their own ceiling. When both exist, the router takes the minimum. That single number becomes the operative ceiling for the request, and candidates whose estimated request cost exceeds it are excluded by the budget filter. The decision records a within-budget reason code whenever a ceiling was in force, so receipts show that cost discipline participated in the outcome.
Note what this mechanism is and is not. It is a per-request eligibility ceiling: a guardrail that prevents routing to candidates the request cannot afford. It is not a cost optimizer, a spend forecast, or evidence of savings. The router makes no claim that the selected candidate is the cheapest acceptable option in any global sense, and nothing in the inspected code measures realized spend against a counterfactual. Teams reading routing receipts should treat the budget marker as what it is — proof a ceiling was enforced — and look elsewhere for any cost analysis. The claim limit on this article exists for exactly this reason: describing a budget gate must never drift into asserting benchmarked savings.
Telemetry: live evidence over static preference
After capability, budget, and health filtering, the router applies live telemetry — recent per-provider, per-model reliability and latency observations. Telemetry enters in two ways: as a hard gate that removes failing candidates, and as scoring input that adjusts the ranking of survivors.
The hard gate is blunt on purpose. A provider and model pair that has failed every recent sample — at least three samples with a zero percent success rate — is treated as unhealthy and excluded with a dedicated telemetry-failing reason, even if it would otherwise win on price or quality. A code comment ties the threshold to the circuit breaker's own tolerance of three consecutive failures: the telemetry gate mirrors that judgment using live Gateway observations as the evidence. The point is structural rather than statistical. Three samples prove little about long-run reliability, but the gate is not estimating reliability — it is refusing to send a new request down a path where every recent attempt just failed.
Surviving candidates carry their telemetry into scoring, where reliability and latency adjustments shift the ranking. The decision records whether telemetry participated, so a receipt can distinguish a telemetry-informed ranking from one decided purely on static signals. Because telemetry is the pipeline's only live input, it is also the only reason two identical requests at different moments can legitimately resolve differently. Determinism holds per snapshot, not across time — a distinction that matters when reproducing a routing decision from logs.
Ranking, scores, and reason codes
Only candidates that survive every gate reach the scorer. Scoring is explainable by construction: each candidate receives a total score plus a breakdown across eight components — quality, latency, time-to-first-token, cost, reliability, health, context, and tools — weighted by the alias's policy. The candidates are ranked by total score, the top rank is selected, and the winner's component breakdown ships with the decision. The exact weights live in alias policy the inspected router consumes but does not define, so this article makes no claim about which component dominates for any particular alias; the mechanism guarantees transparency of the outcome, not uniformity across policies.
The reason codes attached to the decision translate the outcome into a compact, machine-readable story. Every successful decision carries the required-capabilities marker plus any applicable requirement markers (vision, tools, structured output, within-budget, live-certified). The winner's strongest component adds one more code: cost leadership, latency leadership, quality leadership, recent reliability, or a balanced score when no single component dominates. Telemetry-informed decisions add a final marker. Together with the rejected list, these codes let an operator reconstruct why a candidate won without re-running the pipeline — and, just as importantly, let an auditor check that the stated requirements actually constrained the result.
What the ranking must not be mistaken for is a quality verdict on models in general. The winner is the highest-scoring eligible candidate under one alias's weights, for one request's requirements, against one telemetry snapshot. Change the requirements, the alias, or the moment, and the ranking can change. The inspected code contains no cross-task quality benchmark and no "best model" determination; the claim limit prohibiting such assertions reflects what the code actually does. Scoring orders survivors — it does not crown champions.
The canary gate: last check before commitment
After ranking, the decision passes through a canary gate that takes the decision, the request id, and the telemetry map, and may adjust the outcome — including promoting a canary candidate to the selected slot. The router's visible contract is the boundary: the gate observes the full ranked decision and the live telemetry, records its own canary state on the decision, and when it reports a promotion, the selected candidate follows the top rank. The internals of canary selection live outside the inspected router body, so the precise promotion policy is out of scope here; what the router guarantees is that canary handling is explicit, recorded, and applied after — never instead of — the eligibility pipeline. A candidate the filters rejected cannot be canary-promoted back into contention, because promotion only moves within the ranked survivors.
What Chat Faros does instead
Ethen Chat's auto selector, presented in the product as Ethen Faros, is a different mechanism and must not be read as an instance of Gateway smart routing. Where the Gateway filters a catalog pool through gates and scores survivors, Chat resolves auto with a small deterministic policy over thinking depth and tools: deep thinking or a code tool resolves to one explicit large model served through the Vercel AI Gateway, fast thinking resolves to an explicit efficient model on the same transport, and everything else resolves to the default text model served through the direct Meta Model API — never through Gateway entitlement. The Gateway adapter remains in Chat's picture for explicitly selected models, for the deep and fast branches, and for authorized fallback and redundancy, but the normal Faros primary path does not depend on it.
The surrounding rules reinforce the separation. Explicit user selections in Chat never auto-substitute: an unavailable or incapable choice fails closed with an error. Cross-provider fallbacks for auto are fenced behind an enforced opt-out — built only when the caller passes the user's stored safe-fallback grant, which defaults to off — and when enabled they skip remaining models from the failed provider's prefix. Chat's capability inference is conservative by policy: unknown capabilities stay false. Its catalog is fetched live with a ten-second timeout, falls back through an SDK listing to a committed snapshot and finally a small manifest, and caches for five minutes — a resilience ladder for the picker, not a routing score.
The practical upshot: Gateway smart routing answers "which eligible catalog candidate should serve this API request under policy," while Chat Faros answers "which configured default should serve this chat turn given thinking depth and tools." They share vocabulary — capabilities, fallbacks, fail-closed errors — but the code paths, transports, and decision procedures are distinct. Evidence about one is not evidence about the other.
Cortex's role: one signal, not the router
Readers tracing imports will notice the smart router reads circuit-breaker state from the Cortex module. That import is exactly as narrow as it looks: Cortex supplies the open-or-closed circuit signal that feeds provider health and the capability filter's unavailability check. It does not score candidates, merge requirements, build the pool, or make the selection. Naming the boundary precisely matters because "Cortex" appears in several Ethen contexts; in this pipeline it is a health dependency, and routing owns the decision. Collapsing the two into "Cortex routing" would misattribute the pipeline this article describes.
A worked example (synthetic)
Because the inspected sources are code rather than production traces, the following example is synthetic — it illustrates the mechanics without claiming any real request resolved this way. Imagine a request arriving under a smart alias with requirements of tool support and structured output, a route id of text-general, a project id, and a modest budget ceiling.
The router first merges the alias's gates with the request requirements and builds the candidate pool for the text route. Suppose the pool holds a dozen candidates. Provider health resolves each: one provider's circuit is open, so its candidates go unavailable; another lacks an implemented adapter, so its candidates go degraded but survive. The capability filter then removes every candidate without tool support or structured output, every unavailable candidate, any candidate outside the request's provider restriction or the project's allowlist, and any candidate whose estimated cost exceeds the effective ceiling (the lower of the request's and the project's numbers). Suppose five survive.
Telemetry applies next: one survivor's pair shows three recent samples with zero successes, so it is excluded despite strong static scores. The scorer ranks the remaining four under the alias's weights — say the winner leads on reliability. The decision records its reason codes, the canary gate makes no change, and the receipt summarizes the winner, codes, policy version, and counts. Had every candidate been filtered out, the same machinery would have returned a null selection with the no-compatible-candidate code — and no provider call would have followed.
Limitations: what routing does not prove
An honest account of the pipeline ends with what it cannot establish. First, several inputs arrive from modules outside the three inspected sources — alias weights and gate definitions, the capability filter's internals, the scorer's formula, the canary policy, the telemetry collector, the Model Intelligence candidate builder, and the circuit-breaker thresholds — so this article describes their contracts as the router sees them, not their implementations. Behavior inside those boundaries is accurately characterized only at the interface level here.
Second, eligibility is environment-relative. A model that routes today can become unroutable tomorrow through a revoked key, an opened circuit, a changed allowlist, an edited route profile, or a lowered budget — and catalog-only visibility never implied executability in the first place. Routing receipts describe one decision in one environment at one moment; they are not availability promises.
Third, target separation between Ethen's apps is not evidence about any migration. That Gateway routing, Chat Faros, Studio, Research, Designer, Founder, Computer, Desktop local models, and Code's three surfaces are separate concerns is an architecture fact, and this article's routing claims neither depend on nor prove anything about those products' shipping state. Roadmap and implementation stay in separate sentences.
Finally, nothing here supports best-model rankings or cost-saving benchmarks. The router enforces ceilings and orders eligible candidates; it does not run controlled comparisons, and the research program that could produce measured routing comparisons is explicitly future work tracked elsewhere. Any claim that routing "picks the best model" or "saves money" would need evidence this pipeline does not generate — and this article makes neither claim.
The shape to remember
Request, requirements, policy, pool, gates, telemetry, ranking, canary, decision — with rejection recorded at every step and refusal as a valid ending. Ethen Gateway routing is a discipline of saying no for documented reasons until exactly one yes remains, or none does. The next time a receipt lands in your logs, read the rejected list first: the models the Gateway refused, and why, tell you more about your policy than the winner alone ever could.