
# Smart routing (ethen/* aliases)

The Gateway can act as Ethen's **intelligent model execution layer**: instead of
pinning a concrete model id, request an `ethen/*` alias and let the Gateway pick
the provider/model that satisfies the request's hard requirements and best fits
the alias's policy.

Smart routing is **policy-bound, capability-safe, budget-aware, explainable, and
measurable**. It never routes a request to a candidate that cannot satisfy it,
and every decision carries a routing receipt.

## Aliases

| Alias | Behavior | Hard gates |
|---|---|---|
| `ethen/auto` | Balanced quality / cost / latency | none (request requirements still apply) |
| `ethen/fast` | Latency + TTFT dominant | none |
| `ethen/best` | Quality dominant, subject to policy and reliability | none |
| `ethen/cheap` | Cost dominant, subject to minimum quality/reliability | none |
| `ethen/reasoning` | Quality dominant; prefers models with reasoning metadata | none (metadata is a preference, not a gate) |
| `ethen/coding` | Tool-safe coding routing | tools required; context ≥ 64K |
| `ethen/vision` | Vision-capable routing | vision capability required |
| `ethen/long-context` | Large-context routing | context window ≥ 128K |

An alias is only honored when its behavior is fully specified: each alias has a
versioned weight profile (`ethen.routing-policy.v1`) and — where the name
implies a non-negotiable capability — a hard gate. Unknown `ethen/*` ids are
rejected like any unknown model.

## Request

```http
POST /api/gateway/v1/chat/completions
Authorization: Bearer <ethen...key>
```

```json
{
  "model": "ethen/auto",
  "messages": [{ "role": "user", "content": "Summarize this transcript" }],
  "metadata": {
    "routing": {
      "budget_ceiling_usd": 0.05,
      "latency_target": "balanced",
      "certification_level": "any"
    }
  }
}
```

Hard requirements are derived from the request itself:

- **tools** — `tools` array present → tool-capable candidates only
- **vision** — image parts in messages → vision-capable candidates only
- **structured output** — `response_format` present → structured-output-capable candidates only
- **context** — estimated input tokens must fit the candidate window
- **max output** — `max_completion_tokens` / `max_tokens` must fit the candidate
- **budget ceiling** — `metadata.routing.budget_ceiling_usd` (and the project's
  remaining window budget, when configured) excludes candidates whose estimated
  request cost exceeds the ceiling
- **certification** — `metadata.routing.certification_level: "live"` requires a
  current dated live certification receipt
- **provider restrictions** — `providerOptions.gateway.only` narrows the pool

## Routing pipeline

```text
request → requirements → effective project policy → capability filtering
→ Model Intelligence (quality/capability metadata) → live Gateway telemetry
→ budget/cost filter → provider health → ranked candidates → canary gate
→ Gateway execution
```

1. **Requirements extraction** — hard requirements derived from the raw body.
2. **Candidate pool** — built from the Gateway route profiles, enriched by
   Model Intelligence (quality, latency, vision, reasoning, context), the
   pricing registry (authoritative per-token prices), and the model catalog.
3. **Capability filtering** — a candidate that cannot satisfy the request never
   enters the scoring pool (rejection reasons are recorded).
4. **Scoring** — explainable components (quality, latency, TTFT, cost,
   reliability, health, context, tools) weighted by the alias's versioned
   profile. Live telemetry adjusts reliability/latency; providers with a 0%
   success rate over ≥ 3 recent samples are excluded.
5. **Canary gate** — a bounded rollout can promote a target candidate to rank 1
   for a deterministic percentage of requests.
6. **Execution** — the ranked pool becomes the capability-safe fallback chain:
   if the primary fails, only candidates that still satisfy the hard
   requirements are tried.

## Routing receipts

Smart-routed responses include an explainable receipt in the `ethen` namespace
(both buffered responses and the terminal streaming chunk):

```json
{
  "ethen": {
    "routing": {
      "requested_model": "ethen/auto",
      "selected_model": "deepseek-v4-flash",
      "selected_provider": "deepseek",
      "reason_codes": [
        "supports_required_capabilities",
        "within_budget",
        "high_recent_reliability",
        "telemetry_informed"
      ],
      "routing_policy_version": "ethen.routing-policy.v1",
      "candidate_count": 3,
      "rejected_count": 0,
      "score_components": { "quality": 0.4, "cost": 0.85, "latency": 0.6, "ttft": 0.87, "reliability": 0.5, "health": 1, "context": 1, "tools": 0.5 }
    }
  }
}
```

Receipts contain routing facts only — never secrets, telemetry internals, or
tenant identifiers beyond the model/provider selection.

## Telemetry feedback loop

Every executed request reports its real outcome — model, provider, latency,
TTFT, success/failure, error class, tokens, cost, fallback, stream completion —
back into the telemetry store. The router consumes that evidence for reliability
scoring and the failing-provider gate. Telemetry is in-memory and process-local
(it resets on restart); it is operational evidence, never billing.

## Canary routing

Canary targets are operator-configured (empty by default). A target declares an
alias, a candidate, a rollout percent (0 → 1 → 5 → 20 → 50 → 100), and
measurable acceptance criteria (`minSuccessRate`, `minSamples`). Bucketing is
deterministic per (alias, candidate, request id). When the bucket hits, the
canary candidate is promoted to rank 1 for that request; acceptance is measured
from live telemetry and reported on the receipt (`canary.acceptance_met`).

## Behavior notes

- **Determinism** — for a fixed input and telemetry state, the same alias +
  request resolves to the same decision (ties break by provider/model id).
- **Fail-closed** — if the capability filter leaves no pool, the request fails
  with `no_smart_route_candidate` and the rejection reasons; no provider call
  happens.
- **Fallback safety** — the ranked pool is the fallback chain, so a tools +
  vision + large-context request can never fall back to an incompatible model.
- **Optional features** — shadow evaluation and hedged requests are not shipped
  in this milestone; canary routing is the bounded-rollout mechanism.
