Skip to content

EthenEthenEthen

How Ethen Gateway Binds API Requests to Projects

Every chat request carries a key, lands in exactly one project, and leaves a paper trail. Here is how the Gateway enforces that chain.

Every chat request carries a key, lands in exactly one project, and leaves a paper trail. Here is how the Gateway enforces that chain.

Send the Gateway a chat request and three things happen before any model sees your tokens: your key is verified and tied to a project, your model is checked against a runtime catalog, and your request is measured against budgets and idempotency claims. Only then does anything get dispatched to a provider. This article traces that path through the Gateway's chat completions route and its model-catalog status logic, so you know what Ethen gateway API authentication actually guarantees — and where it stops.

The central idea is simple: the key you send inbound is not the credential that pays for inference outbound. Your Bearer key identifies a project and proves you may invoke the Gateway. The provider credential — the server-side key or your project's own stored provider key — is a separate object, resolved later, and never exposed to you. Keeping those two secrets apart is what lets one project hold many keys, rotate them independently, and account every request to the right tenant.

The inbound key: a Bearer token with a project attached

The route reads exactly one credential from the request: the Authorization header, which must have the form Bearer <key>. Anything else — a missing header, a non-Bearer scheme, an empty token — is rejected immediately with a missing_api_key error and HTTP 401, before the body is even parsed for a key lookup. The route does parse the JSON body first to confirm it is valid and that model and messages are present, but no project context exists until the key resolves.

Resolution works in two stages. First the route derives a lookup prefix from the raw key and fetches a bounded set of candidate key records matching that prefix. The prefix exists so the store can narrow candidates without hashing every row. Each candidate's stored hash is then checked against the raw key until one verifies. If none verifies, the caller gets invalid_api_key and a 401 — the same response whether the prefix matched nothing or the hash check failed, so callers cannot probe which part was wrong.

The second stage is where the project binding is enforced. Only keys with a persisted project binding are valid; key records without one are excluded from lookup entirely. In other words, a key that is not attached to exactly one project is not a valid Gateway key, no matter how correctly it was issued otherwise. This is the "binds API requests to projects" guarantee at its most literal — project scope is a property of the key row, assigned at issuance, and every downstream decision (budgets, model eligibility overrides, usage records, stream state) keys off the resolved project_id.

One more detail worth knowing: every authenticated request updates the key's last-used timestamp on a best-effort basis. Rotation audits and stale-key cleanup can rely on it as a signal, though the write is deliberately non-blocking — a logging failure never fails your inference request.

Authorization: scopes, revocation, and expiry

Verifying the hash is authentication; deciding the key may act is authorization, and the Gateway separates the two. After a row verifies, the route runs an authorization check over the key's project binding, scope set, and revocation/expiry state, requiring the gateway:invoke scope.

The outcomes map to distinct errors and status codes:

  • A revoked key is rejected with a revoked-key error and 401.
  • An expired key is rejected with an expired-key error and 401.
  • A valid key without the gateway:invoke scope is rejected with insufficient_scope and 403.

The 401-versus-403 split matters for debugging. A 401 says the credential itself is unusable — check rotation and expiry. A 403 says the credential is fine but lacks permission — the key was issued with a narrower scope set than chat invocation needs. Both failures are still written to the request log (more on that below), under "unknown" when no key record resolved, so rejected attempts are not lost into the void.

What the two inspected sources do not show is how keys are issued, what other scopes exist, or how a caller creates and manages keys. There is no signup, dashboard, or key-management API in these files, and this article makes no claim about them. Treat issuance and rotation tooling as unverified from this evidence, and design-target at best.

Model eligibility: the catalog decides what can run

With the key resolved to a project, the route checks the requested model against the Gateway catalog before spending anything. The check lists catalog models and matches the requested ID against both bare and provider-prefixed forms, so a client can send either style. If nothing matches, the request fails with model_not_found and a 400 — a client error, because the fix is to pick a known model ID, and the error's hint says exactly that.

When the model is known, its catalog status decides the outcome. The status vocabulary has five kinds, and understanding them is the key to reading Gateway errors:

  • runnable means the exact provider/model pair is referenced by the current Gateway route profiles, the provider adapter is configured, and dated certification evidence exists. This is the strictest bar.
  • provider-configured means the provider adapter is configured and the model's capability family is text-compatible, but the pair lacks dated certification evidence or is not explicitly declared in the current route profiles. The route still allows these models through.
  • missing-key means a matching provider exists but the required server-side key is not configured in this environment. This normally fails with provider_key_missing and a 503 — except for one project-scoped escape hatch described below.
  • catalog-only means the model is visible in discovery but the repository has no runtime evidence it can execute. The route rejects it with model_catalog_only and a 400.
  • unsupported-modality means the model's capability family is one the Gateway runtime does not execute — image, video, embedding, rerank, realtime, speech, or transcription. The chat completions route rejects these with a modality error and a 400, since only text-compatible families (text, code, reasoning, long-context) and explicitly handled unknowns can proceed.

Two properties of this design deserve emphasis. First, "runnable" is deliberately hard to earn: it requires the model to appear in default route profiles built from the text route family (text-general, text-quality, text-creative, text-reasoning), a configured adapter, and launch-ready certification. The catalog code even normalizes provider-prefixed IDs against unprefixed route-profile IDs before comparing, because a prefix mismatch once made the runnable state unreachable. Second, the Gateway never silently substitutes a mock when a production provider's key is missing. The status derivation maps missing keys to setup-required regardless of mock-fallback flags — a missing credential surfaces as a configuration error, not a quiet downgrade.

The BYOK exception: your project's own provider key

The missing-key case has one project-aware exception. When the catalog reports missing-key and the request carries a resolved project, the route checks whether the project has its own stored provider credential for that model's provider. If a stored project credential exists for that provider, the model becomes runnable for that request with a synthetic byok-configured status.

This is the cleanest illustration of the inbound/outbound split. The inbound Ethen key proved which project is calling; the outbound provider credential — brought by the project owner, stored server-side, retrieved per provider and project — pays for and authorizes the upstream call. A lookup failure during the BYOK check simply falls through to the standard missing-key rejection; it never leaks whether a credential exists. And note the scoping: BYOK is per project, so two projects requesting the same model can legitimately get different eligibility answers. If you are debugging a provider_key_missing 503, the question is not only "is the provider configured" but "is it configured for this project."

Routing hints are best-effort, never fatal

The request body may carry a providerOptions.gateway namespace with routing controls: a preferred provider order, a provider allow-list, per-provider model aliases, and per-provider timeouts. The route parses these defensively — unknown provider IDs are dropped, malformed values are ignored, and an empty result means "no overrides." An invalid client hint can never 400 a request. The known provider set in the inspected code is mock, openai, anthropic, deepseek, openai-compatible, vercel-ai-gateway, and muse-spark, but treat that as a snapshots of the code, not a supported-provider guarantee: eligibility still depends on catalog status, adapters, and keys as described above. The parsed hints flow into the dispatch layer as fallback order, allow-list, alias overrides, and timeouts; the routing policy itself lives outside the two files inspected here.

Gates before dispatch: budgets, idempotency, and admission

Even with an authorized key and an eligible model, the route passes through three more gates before contacting any provider. Each exists to prevent a different kind of waste or double-spend.

Budgets. The route checks project budget limits with a small estimated cost and an idempotency key (the client's Idempotency-Key header if supplied, otherwise the generated request ID). Enforcement mode comes from GATEWAY_BUDGET_ENFORCEMENT and defaults to soft; the code notes that hard mode reserves atomically via a database procedure while soft mode remains a check-then-act gate. A denial returns budget_exceeded and 429 with a hint naming the window and suggesting a higher limit or a wait. Budgets are project-scoped, which closes the loop on the key-to-project binding: spend control follows the tenant, not the individual key.

Execution idempotency. The budget check doubles as an execution claim. Exactly one caller per (project, idempotency key) pair holds the claim and may dispatch. If the claim is already held, duplicates never reach the provider: a completed execution replays its durable response body with an X-Ethen-Idempotent-Replay header, while an in-progress or terminally failed execution is refused with a 409 and instructions to issue a new key. A second, stream-state-backed idempotency check provides overlapping protection for streaming requests. The design principle is explicit in the comments — reuse of an idempotency key must never re-dispatch billable provider work.

Admission control. When the distributed limiter is enabled, the request must be admitted by a limiter lifecycle bound to the project, actor, request, and key environment. A denial returns 429 (or 503 if the limiter itself is unavailable), and every outcome — allowed, denied, completed, failed, aborted — settles the lifecycle so limiter state stays consistent. When the limiter is disabled, this gate is skipped entirely.

Only after all three gates pass does the route call the dispatch layer with the normalized contract, resolved route ID, project ID, routing hints, and the client's disconnect signal (so upstream calls abort if the client drops mid-stream).

Request accounting: logs, attempts, usage, and receipts

Every request — successful, rejected, or failed mid-stream — leaves structured records. The accounting has four layers, and each answers a different question.

Request logs answer "what happened to this request." Each entry records project ID, key ID, user ID, request and trace IDs, route and model IDs, provider ID, status code, latency, token counts, estimated cost, whether fallback was used, and an error code. Critically, the metadata marks content_logging_mode: metadata_only: the log carries accounting facts, not prompt or completion text. Even authentication failures get a log row (under "unknown" when no key record resolved), so rejection forensics don't require provider-side data.

Provider attempts answer "which upstreams were tried." Each attempt records its sequence number, provider and model IDs, success or failure, error codes, latency, tokens, cost, and price-record reference. On success, only the winning attempt carries token and cost figures. This is the record that makes fallback behavior auditable: the receipt's attempt count is backed by rows, not just a counter.

Usage events answer "what did it cost." Each gateway.request event carries the project, key, user, trace, provider, model, route, token counts, estimated cost, and a usageEstimated flag that is set whenever exact provider token data was unavailable — for example, when output tokens had to be estimated from text length, or when an error path had no usage data at all. Usage events are written even on error paths, including failed streams (with partial token estimates) and terminal failures, so cost visibility survives unhappy paths.

Receipts and identifiers answer "how do I correlate all of this." Every response carries an ethen object with the request ID, trace ID, project ID, API key ID, upstream provider, and attempt count. Non-streaming responses use chatcmpl_-prefixed IDs with a chat.completion shape; streaming responses relay chat.completion.chunk events over server-sent events with per-chunk sequence numbers, character offsets, time-to-first-token, and a terminal [DONE] sentinel, plus X-Gateway-Request-Id, X-Gateway-Trace-Id, and X-Gateway-Stream-Id headers. Both paths also return an X-Ethen-Model-Status header echoing the catalog status (runnable, provider-configured, byok-configured, and so on) so clients can see which eligibility path their request took.

Stream persistence follows the same metadata-only discipline: stream state, segments, and events record offsets, hashes, and payloads of control events — never raw content. Heartbeats keep long streams alive, and a mid-stream provider failure produces a structured error event rather than a misleading clean finish.

Errors: whose fault, and what to do

The route's error taxonomy is worth reading as a debugging guide. Client errors (4xx) mean you can fix the request: missing or invalid keys (401), missing scope (403), unknown or non-runnable models (400), budget exhaustion (429), idempotency conflicts (409). Server errors (5xx) mean the problem is upstream or environmental: missing provider keys and unavailable catalogs (503), unavailable limiter (503). One deliberate remap sits between the Gateway and its providers: upstream 401s, 403s, and other provider 4xx responses are converted to 502s (upstream_auth_error, upstream_provider_error), because a provider rejecting the Gateway's credential is a Gateway-side configuration problem, not a signal that your request was malformed. Provider 503s and timeouts pass through unchanged.

Every error carries a trace ID, and most carry a docs_hint with a concrete next step. When reporting an issue, that trace ID plus the request ID is what connects your report to the project's log rows.

Limitations: what this evidence does not establish

These two files describe the request path precisely, but their scope ends at clear boundaries, and staying honest about them matters.

First, nothing here describes key issuance, rotation workflows, or any management surface. Scopes beyond gateway:invoke are not enumerated, and no claim is made about how a developer obtains a key.

Second, nothing here is a launch or availability statement. The Gateway route's existence in the repository does not prove a public API is live, priced, or reachable at any particular host — and this article names no host for exactly that reason. Availability claims belong to a launch announcement with deployment evidence, not to a code walkthrough.

Third, response shapes resembling a well-known chat API are observed behavior of this route, not a compatibility guarantee. Client libraries, streaming edge cases, and tool-call handling may differ; verify against the actual route before assuming drop-in behavior.

Fourth, cost figures are estimates. Token counts are sometimes estimated from text, usage events flag when they are approximate, and pricing depends on records outside these files. Treat Gateway cost data as operational telemetry, not invoicing.

Finally, the provider set, route profiles, and certification states are snapshots of the inspected code. What is runnable today depends on live configuration — adapters, keys, route profiles, and dated evidence — not on any list printed here.

Tying it together

A Gateway request is a chain of bindings: a Bearer key resolves to exactly one project; the project passes budget and idempotency gates; the model passes catalog eligibility, possibly via the project's own provider credential; and every outcome lands in project-scoped logs, attempts, and usage events stitched together by request and trace IDs. Inbound identity and outbound payment stay separate at every step, which is what makes per-project accounting possible at all. If you take one debugging habit from this article, let it be this: when a request fails, read the error code for whose problem it is, then follow the trace ID — the Gateway wrote down everything it decided, and why.