Skip to content

EthenEthenEthen

How Ethen Desktop Talks to Local Models

Ethen Desktop keeps every local model operation behind nine narrow IPC channels and a provider capability matrix, so the UI can only ask for what the runtime actually supports.

Ethen Desktop keeps every local model operation behind nine narrow IPC channels and a provider capability matrix, so the UI can only ask for what the runtime actually supports.

The Ethen Desktop local runtime does not give the window a network client. It gives it a menu: nine named operations, one event channel for progress, and a capability record that says which operations each provider kind supports. Everything privileged — reaching the local endpoint, downloading a model, deleting one — stays in the main process. The renderer asks; the main process decides.

That shape matters because local runtimes are not interchangeable. The current code treats Ollama as the fully supported path, with its own /api/* endpoints, while LM Studio, llama.cpp, and custom OpenAI-compatible endpoints share a narrower /v1-style surface: they can report status, list models, and stream chat, but they cannot install or delete models through Ethen. This article walks through both halves of that design as implemented: the capability matrix in runtime-capabilities.ts and the IPC layer in local-models.ts.

Why the main process owns the runtime

The header comment in the Desktop IPC module states the ownership rule directly: the main process owns all privileged access to the local runtime endpoint, and the renderer can only call named, allowlisted methods through the bridge. The module then lists the security rules that follow from that decision: only localhost endpoints are allowed, model names are validated before any adapter call, no arbitrary URL fetch is exposed to the renderer, no raw ipcRenderer, Node, fs, child_process, or shell access is exposed, no secrets live in renderer state, pull and chat progress stream through a single allowlisted event channel, and a request registry supports cancellation without exposing a raw AbortController.

This is a deliberately narrow waist. If the renderer could fetch arbitrary URLs or hold credentials, a compromised window could pivot anywhere. Instead, the renderer's vocabulary is fixed to the nine channels in the CHANNELS constant: local-models:status, local-models:list-installed, local-models:list-running, local-models:list-catalog, local-models:show, local-models:pull, local-models:cancel, local-models:delete, and local-models:chat. Anything else is not a degraded operation; it is simply not reachable.

Desktop owns Local Models in Ethen's product map, which is why this enforcement lives here rather than in Chat or Platform code. Chat stays lightweight, and Computer remains Platform-only. Code spans all three surfaces — lightweight Chat help, cloud Platform workspaces, and the local Desktop environment — so a developer working in Code locally may end up exercising this same Desktop path. But the IPC boundary described here is a Desktop implementation detail, not a claim about what any other app exposes or what has shipped as a public download.

The capability matrix: one record per provider kind

The second file, runtime-capabilities.ts, answers a different question: given a provider kind, which operations should the UI even offer? Its grounding rule is conservative by default — if a provider kind is unknown or its capability is not yet implemented, the answer is false.

There are five provider kinds: ollama, lm-studio, llama-cpp, custom-openai, and unknown. Each maps to a LocalRuntimeCapabilities record with twelve fields: canCheckStatus, canListInstalledModels, canListRunningModels, canShowModelDetails, canPullModels, canDeleteModels, supportsChatStreaming, supportsEmbeddings, optional supportsTools, supportsVision, supportsJsonMode, and openAiCompatible.

The table below restates the registry as written. Read the LATER suffixes on three of the four non-Ollama records as what they are: labels in the source indicating those records describe a known API surface, not necessarily a wired-up integration. Target separation in the code is not proof that any of those paths has shipped.

CapabilityOllama (MVP)LM Studio (later)llama.cpp (later)Custom OpenAI (later)Unknown
Check statusYesYes (/v1/models health check)Yes (/v1/models health check)Yes (/v1/models or health)No
List installedYesYes (/v1/models)Yes (best-effort)Yes (/v1/models)No
List runningYesNo (no running-vs-installed distinction)No (single-model server)NoNo
Show detailsYesNo (no detail endpoint)NoNoNo
Pull / installYesNo (managed in its own GUI)No (does not install)NoNo
DeleteYesNo (no delete API)NoNoNo
Chat streamingYesYesYes (/v1/chat/completions SSE)YesNo
EmbeddingsNoNoNoNoNo
Tools / vision / JSON modeNoNoNoNoNo
OpenAI-compatibleNo (own /api/*)YesYesYesNo

Two contrasts carry most of the practical weight. First, only Ollama supports install, delete, running-model listing, and detail lookup. Every OpenAI-compatible record supports status, installed-model listing, and chat streaming, and nothing else. Second, Ollama is explicitly marked as not OpenAI-compatible because it uses its own /api/* endpoints rather than /v1. There is no universal install path across runtimes, and the code refuses to pretend otherwise: unknown providers get a record where every field is false.

Gating UI actions without hardcoding provider names

Components are not supposed to branch on strings like "lm-studio". The module provides lookup helpers so the UI gates actions through the matrix instead. getRuntimeCapabilities returns the full record for a kind, falling back to the minimal unknown record. Three focused helpers — canInstallModels, canDeleteModels, and canChatStream — answer the questions buttons most often need.

The most developer-useful helper is getUnsupportedActionReason. It takes a provider kind and one of eight action names (status, listInstalled, listRunning, show, pull, delete, chat, embeddings) and returns null when the action is supported or a deterministic, UI-safe reason string when it is not. The generic form reads <provider> does not support <action label>, using labels like "install models" for pull and "stream chat completions" for chat. But three cases get specific guidance:

  • Pulling on LM Studio or llama.cpp explains that the provider does not support installing models and points at that program's own GUI for model management.
  • Deleting on those two providers gives the parallel message for deletion.
  • Pulling or deleting on a custom OpenAI-compatible endpoint states plainly that custom endpoints do not support installation or deletion.

Unknown actions produce an Unknown action message rather than silently returning false, which keeps typos visible during development. Because the strings are deterministic, a component can render them directly or snapshot them in tests without worrying about changing copy.

A typical gating pattern therefore looks like this: before rendering an Install button, call canInstallModels(kind); when it returns false, call getUnsupportedActionReason(kind, "pull") and show the result as the disabled reason. The matrix stays the single source of truth, and adding a future provider means adding one record, not hunting through components.

Nine channels, each with one job

Back in the IPC module, registerLocalModelsIpc wires each channel to exactly one handler. The six read-oriented operations are straightforward request/response pairs:

  • Status delegates to getOllamaStatus and returns the runtime health record.
  • List installed delegates to listOllamaModels and returns status plus the model list.
  • List running calls safeOllamaFetch("/api/ps") and returns the models array, defaulting to an empty list when the shape is unexpected.
  • List catalog is the honest outlier: it always returns an empty list with a limitation string explaining that Ollama provides no searchable catalog API, and pointing the user at ollama search <query> on the CLI or the Ollama search page in a browser. The channel exists so the UI has somewhere to ask; the answer is that there is nothing to ask.
  • Show validates the model name, then delegates to showOllamaModel for per-model metadata.
  • Cancel takes a requestId string, looks it up in the active-request registry, marks it cancelled, and aborts its controller. Unknown or already-finished IDs get an explicit error rather than a silent success.

The implementation behind status, installed-model listing, detail lookup, and pull streaming lives in the infrastructure layer (apps/infrastructure/lib/local-models), with shared types imported from lib/local-models/types.ts. The Desktop module is the policy and transport shell around those calls: validation, approval, timeouts, and event routing happen here.

Pull and chat stream through one event channel

The two long-running operations — pull and chat — share a pattern. The ipcMain.handle callback validates the request, registers an AbortController in the activeRequests map under a fresh UUID, and returns { ok: true, requestId } immediately. The real work then runs in a background async block that pushes typed events to the renderer over the single local-models:event channel.

For pull, the background task calls streamLocalModelPull with the validated model name, the registered abort signal, and an onProgress callback that forwards each progress event as { type: "pull:progress", requestId, event }. Completion, cancellation, and failure arrive as pull:complete, pull:cancelled, or pull:error, each carrying the model name and, for the failure cases, an error string. The cancelled-versus-error distinction comes from the registry entry's cancelled flag, so a user-initiated abort is reported as a cancellation rather than a crash.

Chat follows the same shape with its own event trio: chat:delta for each streamed content fragment, chat:complete at the end, and chat:error on failure. The request must include a model string and a messages array, and every message must carry both a role and content; anything else is rejected before a request ID is issued. The streaming helper posts { model, messages, stream: true } to /api/chat, reads the newline-delimited JSON response line by line, forwards each message's content as a delta, and treats a done: true line as completion. Unparseable lines are skipped rather than failing the whole stream, and a two-minute overall timeout bounds the call.

This design keeps two invariants intact. The renderer never holds the abort controller — cancellation goes through the cancel channel with an opaque ID. And progress never arrives on ad-hoc channels — every event, for both operations, flows through the one allowlisted name, which is exactly what the bridge can expose without opening new surface.

Mutations need the main process to ask a human

The section of the module marked LM-P0-03 handles the most security-sensitive choice: who approves a download or deletion. The answer is the main process, by asking the desktop user directly through a native confirmation dialog. The renderer may request a mutation, but any approval object it supplies is explicitly ignored — the code comments note that a compromised renderer could trivially forge { approved: true }.

The default dialog uses Electron's native message box with the warning style. A pull shows "Download \<model\>?" with the acknowledgement "I understand this will download the model from the public Ollama registry to my local disk using my network." A delete shows "Delete \<model\>?" with "I understand this will permanently remove the selected local model from my disk." Both offer Cancel as the safe default alongside the destructive action, and only an explicit confirmation lets the handler proceed. Declining produces a plain error: the download or delete was cancelled because the desktop user did not confirm.

The dialog function is injectable through the confirmMutation option on registerLocalModelsIpc, so tests can assert the approval behavior without opening a real window. Delete then issues a DELETE to /api/delete with the validated model name; pull starts the streaming flow described above. There is no silent-install path anywhere in this module, and no batch or one-click operation that skips the dialog. Each mutation is one model, one prompt, one decision.

Validation at every boundary

Two validators run before any runtime call. Model names pass through assertOllamaModelName from the shared runtime core, which rejects empty strings, overlong names, path traversal, whitespace, control characters, and URL-like content. The IPC module keeps its own earlier regex and checks as commented-out reference, but the live path delegates to the shared assertion so Desktop and any other consumer enforce the same rule.

The endpoint itself resolves through resolveOllamaEndpoint, and the module's commented reference implementation documents the intended localhost-only policy: http only, no embedded credentials, host restricted to localhost, 127.0.0.1, or ::1, and port fixed at 11434. The live resolveBaseUrl delegates to the shared resolver. Either way, the architectural point stands: the renderer never supplies a host, port, or scheme. The only address the main process will talk to is the local Ollama endpoint, which is what makes "local" a structural guarantee rather than a configuration suggestion.

The fetch helper adds the remaining guardrails: redirects are treated as errors, a 15-second timeout bounds metadata calls, non-2xx responses become explicit HTTP-status errors, and malformed JSON surfaces as an "unexpected response" error instead of propagating half-parsed data. On shutdown, cancelAllLocalModelRequests aborts every outstanding controller and clears the registry.

What Desktop explicitly does not do

The limitations are as informative as the features, and the code states most of them outright:

  • No catalog search. The list-catalog handler documents that Ollama has no search API and redirects to the CLI or website. Any UI built on this channel should present that guidance, not an empty state that looks like a bug.
  • No embeddings, tools, vision, or JSON mode. Every capability record sets supportsEmbeddings to false, and all three optional flags default to false. Chat streaming is the only generation operation the matrix acknowledges.
  • No install or delete outside Ollama. The three OpenAI-compatible records all set canPullModels and canDeleteModels to false, with reason strings that direct users to each program's own management UI. There is no universal one-click install in this design, and adding one would contradict the matrix rather than extend it.
  • No remote endpoints. Localhost-only resolution plus redirect: "error" means the IPC layer cannot be repurposed as a generic HTTP client, even by a caller that controls the model name.
  • No renderer-held privilege. No secrets, no raw controllers, no Node APIs cross the bridge. Cancellation and progress work through IDs and events precisely so the window stays unprivileged.

These boundaries also scope what this article can claim. The code shows an IPC and capability design; it does not show a public Desktop binary, an installer, an auto-update channel, or benchmark results for local inference. This publication does not cover launch or availability — those questions belong to a release-gated story with its own evidence, not to a code walkthrough. Likewise, the LATER records for LM Studio, llama.cpp, and custom endpoints describe known API surfaces for future gating; they are not evidence that those integrations are wired, tested, or shipped.

Putting it together: a developer's mental model

If you are building against this layer, the mental model has three steps. First, ask the matrix: call getRuntimeCapabilities for the active provider kind and gate every button on its record, using getUnsupportedActionReason for disabled-state copy. Second, call narrow channels: status, list, show, pull, cancel, delete, and chat each do one thing, validate their inputs, and return either a typed success or { ok: false, error }. Third, follow request IDs for anything long-running: pull and chat return immediately with an ID, stream typed events on local-models:event, and cancel through the cancel channel.

The pull flow shows all three steps at once. The UI checks canInstallModels("ollama"), invokes local-models:pull with a model name, receives a request ID, renders pull:progress events, and offers a cancel button that invokes local-models:cancel with that ID — while the main process handles endpoint resolution, name validation, the native approval dialog, streaming, and cleanup. The chat flow is the same shape with deltas instead of byte counts.

That is the whole contract: a small vocabulary, explicit capability differences between Ollama and OpenAI-compatible runtimes, and human approval standing between the window and the disk. Everything else — catalog browsing, model management for other runtimes, richer generation modes — is either delegated to the right tool or honestly marked unsupported.