Smart routing (ethen/* aliases)
GW-R6 — intelligent model execution: ethen/* smart aliases, capability-safe filtering, explainable scoring, routing receipts, live telemetry, and canary routing.
The Gateway can act as Ethen's intelligent model execution layer: instead of pinning a concrete model id, request an ethen/* alias and let the Gateway pick the provider/model that satisfies the request's hard requirements and best fits the alias's policy.
Smart routing is policy-bound, capability-safe, budget-aware, explainable, and measurable. It never routes a request to a candidate that cannot satisfy it, and every decision carries a routing receipt.
Aliases
| Alias | Behavior | Hard gates |
|---|---|---|
ethen/auto | Balanced quality / cost / latency | none (request requirements still apply) |
ethen/fast | Latency + TTFT dominant | none |
ethen/best | Quality dominant, subject to policy and reliability | none |
ethen/cheap | Cost dominant, subject to minimum quality/reliability | none |
ethen/reasoning | Quality dominant; prefers models with reasoning metadata | none (metadata is a preference, not a gate) |
ethen/coding | Tool-safe coding routing | tools required; context ≥ 64K |
ethen/vision | Vision-capable routing | vision capability required |
ethen/long-context | Large-context routing | context window ≥ 128K |
An alias is only honored when its behavior is fully specified: each alias has a versioned weight profile (ethen.routing-policy.v1) and — where the name implies a non-negotiable capability — a hard gate. Unknown ethen/* ids are rejected like any unknown model.
Request
POST /api/gateway/v1/chat/completions
Authorization: Bearer <ethen...key>Hard requirements are derived from the request itself:
- tools —
toolsarray present → tool-capable candidates only - vision — image parts in messages → vision-capable candidates only
- structured output —
response_formatpresent → structured-output-capable candidates only - context — estimated input tokens must fit the candidate window
- max output —
max_completion_tokens/max_tokensmust fit the candidate - budget ceiling —
metadata.routing.budget_ceiling_usd(and the project's
remaining window budget, when configured) excludes candidates whose estimated request cost exceeds the ceiling
- certification —
metadata.routing.certification_level: "live"requires a
current dated live certification receipt
- provider restrictions —
providerOptions.gateway.onlynarrows the pool
Routing pipeline
request → requirements → effective project policy → capability filtering
→ Model Intelligence (quality/capability metadata) → live Gateway telemetry
→ budget/cost filter → provider health → ranked candidates → canary gate
→ Gateway execution- Requirements extraction — hard requirements derived from the raw body.
- Candidate pool — built from the Gateway route profiles, enriched by
Model Intelligence (quality, latency, vision, reasoning, context), the pricing registry (authoritative per-token prices), and the model catalog.
- Capability filtering — a candidate that cannot satisfy the request never
enters the scoring pool (rejection reasons are recorded).
- Scoring — explainable components (quality, latency, TTFT, cost,
reliability, health, context, tools) weighted by the alias's versioned profile. Live telemetry adjusts reliability/latency; providers with a 0% success rate over ≥ 3 recent samples are excluded.
- Canary gate — a bounded rollout can promote a target candidate to rank 1
for a deterministic percentage of requests.
- Execution — the ranked pool becomes the capability-safe fallback chain:
if the primary fails, only candidates that still satisfy the hard requirements are tried.
Routing receipts
Smart-routed responses include an explainable receipt in the ethen namespace (both buffered responses and the terminal streaming chunk):
{
"ethen": {
"routing": {
"requested_model": "ethen/auto",
"selected_model": "deepseek-v4-flash",
"selected_provider": "deepseek",
"reason_codes": [
"supports_required_capabilities",
"within_budget",
"high_recent_reliability",
"telemetry_informed"
],
"routing_policy_version": "ethen.routing-policy.v1",
"candidate_count": 3,
"rejected_count": 0,
"score_components": { "quality": 0.4, "cost": 0.85, "latency": 0.6, "ttft": 0.87, "reliability": 0.5, "health": 1, "context": 1, "tools": 0.5 }
}
}
}Receipts contain routing facts only — never secrets, telemetry internals, or tenant identifiers beyond the model/provider selection.
Telemetry feedback loop
Every executed request reports its real outcome — model, provider, latency, TTFT, success/failure, error class, tokens, cost, fallback, stream completion — back into the telemetry store. The router consumes that evidence for reliability scoring and the failing-provider gate. Telemetry is in-memory and process-local (it resets on restart); it is operational evidence, never billing.
Canary routing
Canary targets are operator-configured (empty by default). A target declares an alias, a candidate, a rollout percent (0 → 1 → 5 → 20 → 50 → 100), and measurable acceptance criteria (minSuccessRate, minSamples). Bucketing is deterministic per (alias, candidate, request id). When the bucket hits, the canary candidate is promoted to rank 1 for that request; acceptance is measured from live telemetry and reported on the receipt (canary.acceptance_met).
Behavior notes
- Determinism — for a fixed input and telemetry state, the same alias +
request resolves to the same decision (ties break by provider/model id).
- Fail-closed — if the capability filter leaves no pool, the request fails
with no_smart_route_candidate and the rejection reasons; no provider call happens.
- Fallback safety — the ranked pool is the fallback chain, so a tools +
vision + large-context request can never fall back to an incompatible model.
- Optional features — shadow evaluation and hedged requests are not shipped
in this milestone; canary routing is the bounded-rollout mechanism.