Beta

Smart routing (ethen/* aliases)

GW-R6 — intelligent model execution: ethen/* smart aliases, capability-safe filtering, explainable scoring, routing receipts, live telemetry, and canary routing.

Raw

The Gateway can act as Ethen's intelligent model execution layer: instead of pinning a concrete model id, request an ethen/* alias and let the Gateway pick the provider/model that satisfies the request's hard requirements and best fits the alias's policy.

Smart routing is policy-bound, capability-safe, budget-aware, explainable, and measurable. It never routes a request to a candidate that cannot satisfy it, and every decision carries a routing receipt.

Aliases

AliasBehaviorHard gates
ethen/autoBalanced quality / cost / latencynone (request requirements still apply)
ethen/fastLatency + TTFT dominantnone
ethen/bestQuality dominant, subject to policy and reliabilitynone
ethen/cheapCost dominant, subject to minimum quality/reliabilitynone
ethen/reasoningQuality dominant; prefers models with reasoning metadatanone (metadata is a preference, not a gate)
ethen/codingTool-safe coding routingtools required; context ≥ 64K
ethen/visionVision-capable routingvision capability required
ethen/long-contextLarge-context routingcontext window ≥ 128K

An alias is only honored when its behavior is fully specified: each alias has a versioned weight profile (ethen.routing-policy.v1) and — where the name implies a non-negotiable capability — a hard gate. Unknown ethen/* ids are rejected like any unknown model.

Request

POST /api/gateway/v1/chat/completions
Authorization: Bearer <ethen...key>

Hard requirements are derived from the request itself:

  • tools — tools array present → tool-capable candidates only
  • vision — image parts in messages → vision-capable candidates only
  • structured output — response_format present → structured-output-capable candidates only
  • context — estimated input tokens must fit the candidate window
  • max output — max_completion_tokens / max_tokens must fit the candidate
  • budget ceiling — metadata.routing.budget_ceiling_usd (and the project's

remaining window budget, when configured) excludes candidates whose estimated request cost exceeds the ceiling

  • certification — metadata.routing.certification_level: "live" requires a

current dated live certification receipt

  • provider restrictions — providerOptions.gateway.only narrows the pool

Routing pipeline

text
request → requirements → effective project policy → capability filtering
→ Model Intelligence (quality/capability metadata) → live Gateway telemetry
→ budget/cost filter → provider health → ranked candidates → canary gate
→ Gateway execution
  1. Requirements extraction — hard requirements derived from the raw body.
  2. Candidate pool — built from the Gateway route profiles, enriched by

Model Intelligence (quality, latency, vision, reasoning, context), the pricing registry (authoritative per-token prices), and the model catalog.

  1. Capability filtering — a candidate that cannot satisfy the request never

enters the scoring pool (rejection reasons are recorded).

  1. Scoring — explainable components (quality, latency, TTFT, cost,

reliability, health, context, tools) weighted by the alias's versioned profile. Live telemetry adjusts reliability/latency; providers with a 0% success rate over ≥ 3 recent samples are excluded.

  1. Canary gate — a bounded rollout can promote a target candidate to rank 1

for a deterministic percentage of requests.

  1. Execution — the ranked pool becomes the capability-safe fallback chain:

if the primary fails, only candidates that still satisfy the hard requirements are tried.

Routing receipts

Smart-routed responses include an explainable receipt in the ethen namespace (both buffered responses and the terminal streaming chunk):

json
{
  "ethen": {
    "routing": {
      "requested_model": "ethen/auto",
      "selected_model": "deepseek-v4-flash",
      "selected_provider": "deepseek",
      "reason_codes": [
        "supports_required_capabilities",
        "within_budget",
        "high_recent_reliability",
        "telemetry_informed"
      ],
      "routing_policy_version": "ethen.routing-policy.v1",
      "candidate_count": 3,
      "rejected_count": 0,
      "score_components": { "quality": 0.4, "cost": 0.85, "latency": 0.6, "ttft": 0.87, "reliability": 0.5, "health": 1, "context": 1, "tools": 0.5 }
    }
  }
}

Receipts contain routing facts only — never secrets, telemetry internals, or tenant identifiers beyond the model/provider selection.

Telemetry feedback loop

Every executed request reports its real outcome — model, provider, latency, TTFT, success/failure, error class, tokens, cost, fallback, stream completion — back into the telemetry store. The router consumes that evidence for reliability scoring and the failing-provider gate. Telemetry is in-memory and process-local (it resets on restart); it is operational evidence, never billing.

Canary routing

Canary targets are operator-configured (empty by default). A target declares an alias, a candidate, a rollout percent (0 → 1 → 5 → 20 → 50 → 100), and measurable acceptance criteria (minSuccessRate, minSamples). Bucketing is deterministic per (alias, candidate, request id). When the bucket hits, the canary candidate is promoted to rank 1 for that request; acceptance is measured from live telemetry and reported on the receipt (canary.acceptance_met).

Behavior notes

  • Determinism — for a fixed input and telemetry state, the same alias +

request resolves to the same decision (ties break by provider/model id).

  • Fail-closed — if the capability filter leaves no pool, the request fails

with no_smart_route_candidate and the rejection reasons; no provider call happens.

  • Fallback safety — the ranked pool is the fallback chain, so a tools +

vision + large-context request can never fall back to an incompatible model.

  • Optional features — shadow evaluation and hedged requests are not shipped

in this milestone; canary routing is the bounded-rollout mechanism.

Last verified 2026-08-08 · Owner gateway-team