Local AI or GPU Hosting: What to Check First
A capability-first checklist for deciding where an open-weight model runs — local runtime support on one side, deployment recipes on the other.
A capability-first checklist for deciding where an open-weight model runs — local runtime support on one side, deployment recipes on the other.
The local AI vs GPU hosting question usually starts with hardware. A better starting question is operational: which actions do you need to perform — install a model, list what is running, stream a chat, verify a deployment — and which execution path actually supports each one? Hardware matters later. Capability support decides first, because an execution path that cannot perform your required operation is disqualified no matter how attractive the rest of the setup looks.
This guide gives you a durable, capability-first checklist. It is grounded in three inspected implementation sources: a local runtime capability matrix, a desktop IPC layer for local models, and a versioned GPU deployment recipe registry. The specific code will evolve, but the checks transfer: every local runtime and every hosted path answers the same questions about supported operations, management boundaries, health verification, and versioning. Work through them before you commit.
Start with operations, not machines
List the operations your workflow needs before comparing paths. The inspected capability matrix names eight explicitly: check runtime status, list installed models, list running models, show model details, install (pull) models, delete models, stream chat completions, and generate embeddings. That list is a useful template because each entry is independently supported or unsupported depending on the provider. "Supports local models" is not one claim; it is eight separate answers.
Write your own required-operation list in the same style. A drafting workflow might need only status checks and streaming chat. A model experimentation workflow needs install, delete, detail inspection, and the ability to distinguish installed models from currently running ones. A deployment workflow needs health verification and repeatable provisioning instead. Once the list exists, each candidate path either supports an entry or it does not — and "does not" is a normal, expected outcome to design around, not a failure.
One more scoping rule applies throughout this guide. The Ethen architecture used as reference keeps product surfaces separate: Desktop owns Local Models, while heavier execution lives elsewhere, and each target app has its own boundary. Treat that separation as a design target described by the batch authority, not as proof that any particular migration or launch has shipped. The checks below concern what each path supports, not what is publicly available.
Check 1: Does the local runtime support each operation?
Local runtimes differ more than their marketing suggests. The inspected capability matrix covers four provider kinds — Ollama, LM Studio, llama.cpp, and custom OpenAI-compatible endpoints — plus an unknown kind that answers "no" to everything by conservative default. Unknown or unwired providers returning false is deliberate grounding behavior: the UI gates actions on these flags rather than hardcoding provider names, so an unrecognized runtime disables operations instead of attempting ones it cannot perform.
The differences are concrete. In the inspected snapshot, only the Ollama record supports installing and deleting models through its API, listing running models separately from installed ones, and showing per-model details. LM Studio, llama.cpp, and custom endpoints support status checks, installed-model listing, and streaming chat, but not running-model listing, detail views, install, or delete. Embeddings read false across all four records in this snapshot, as do tool calling, vision input, and JSON mode. Streaming chat is the one capability every known record shares.
Two durable lessons follow. First, verify operations individually against whatever matrix your tooling exposes; never infer the full set from one working feature. Chat streaming working says nothing about install, delete, or embeddings. Second, treat conservative defaults as a feature. A runtime layer that disables unsupported actions with an explicit reason is easier to reason about than one that lets you attempt them and fail ambiguously.
Check 2: Which endpoint dialect does your tooling speak?
A runtime can support an operation while speaking a different protocol than your client expects. The inspected matrix tracks this with an explicit openAiCompatible flag: LM Studio, llama.cpp, and custom endpoints expose an OpenAI-compatible /v1 surface for chat, while Ollama uses its own /api/* endpoints and is flagged not OpenAI-compatible. Same operation — streaming chat — two dialects.
This matters because client code, SDKs, and gateway integrations are usually written against one dialect. If your tooling posts to /v1/chat/completions with server-sent-event streaming, it aligns with the three OpenAI-compatible records and needs an adapter or a separate code path for the native /api/chat shape. The inspected desktop code demonstrates the point: its chat helper posts to /api/chat and parses newline-delimited JSON lines with per-line message.content extraction, which is Ollama-native handling, not /v1 handling.
Your checklist entry is therefore a question, not an assumption: for each runtime candidate, which endpoint shape does chat use, and does every client in my workflow speak it? Keep this check separate from Check 1. Capability ("supports streaming chat") and protocol ("over which endpoint shape") are independent axes, and both must match.
Check 3: Who manages the models — your app or another GUI?
Install and delete are the operations most likely to differ between local paths, and the reason is architectural rather than incidental. In the inspected matrix, only Ollama exposes model installation and deletion through its API. LM Studio manages models through its own GUI and exposes no install or delete API; llama.cpp, as a single-model server, has no install concept at all; custom OpenAI-compatible endpoints likewise do not install or delete.
The implementation goes further by encoding this in user-facing guidance. When an unsupported install or delete is attempted, the reason strings direct the user to the owning GUI — LM Studio or llama.cpp — or state plainly that custom endpoints do not support installation. This is the durable pattern to check for in any local setup: model lifecycle ownership must live exactly one place, and every other surface should say so explicitly instead of offering a broken button.
Ask of each candidate path: where do models get installed and removed, and what happens when code asks a non-owning surface to do it? Acceptable answers include "the runtime API handles it," "the vendor GUI handles it and the API says so," or "nothing handles it yet and the UI disables the action." The unacceptable answer is a control that appears to work and fails at runtime. If you need in-app install and delete, that requirement alone narrows the field to runtimes whose APIs actually expose both.
Check 4: What does the application boundary allow?
Even when the runtime supports an operation, the application sitting in front of it may narrow what is reachable — and that narrowing is often intentional. The inspected desktop IPC layer is a clear example of a deliberately narrow boundary, and its rules make a good checklist for evaluating any local-models frontend.
First, endpoints are localhost-only. The main process owns all privileged access to the runtime endpoint, restricted to local loopback hosts on port 11434, and the renderer cannot fetch arbitrary URLs. Second, model names are validated before any adapter call, rejecting empty strings, overlong input, path traversal patterns, whitespace, control characters, and URL-like structure. Third, mutations require main-process confirmation: pull and delete proceed only after the desktop user confirms in a native dialog, and any approval object supplied by the renderer is ignored, because a compromised renderer could forge one. The dialog text states the consequence plainly — a pull downloads the model from the public registry to local disk over the user's network; a delete permanently removes it from disk. Fourth, long operations are cancellable through a request registry without exposing raw abort controllers to the renderer, and all requests are cleaned up at shutdown. Fifth, progress and chat deltas stream back over a single allowlisted event channel rather than ad-hoc callbacks.
Two limitations in the same file are equally instructive. The catalog operation reports that Ollama provides no searchable model catalog API and returns an empty list with guidance to use the CLI or browse the registry website instead. Chat requests require a model plus messages with roles and content, carry a two-minute streaming window, and ordinary fetches time out after fifteen seconds. An honest boundary documents its timeouts and its missing operations.
Your checklist version: which channels does the app expose, what validation runs before the runtime is touched, who confirms destructive actions and where, how are long operations cancelled, and which operations return explicit "not provided" answers? A boundary that answers all five is one you can build on. Also note what this checklist deliberately excludes: it says nothing about speed, power draw, privacy properties, or cost. Those need fresh, scoped measurements for your hardware and workload, and no capability matrix supplies them.
Check 5: Is there a versioned recipe for the hosted path?
The hosted side of the decision replaces "which runtime operations exist" with "which deployment recipes exist and what each one guarantees." The inspected recipe registry models this well, and its structure transfers to any GPU hosting evaluation.
Each recipe is a versioned, immutable record selected as the template field at instance creation time, with no per-launch remote command execution. Immutability is enforced: registering a duplicate id-plus-version throws, and stored records are frozen. Retrieval supports fetching a specific version, resolving the latest version of an id, listing one record per id, and filtering to enabled recipes. The registry therefore answers the two questions every deployment path must answer: "which exact version am I deploying?" and "can that version change underneath me?" Here the answers are "a pinned id-plus-version" and "no."
The built-in records show what a complete recipe declares: a stable id and human-readable name, a semantic version, a description, a backing template name, a kind distinguishing native provider templates from Ethen golden snapshots, expected open ports, a health check specification, an enabled flag, a provenance string, a minimum disk size, an optional recommended GPU type, and an immutable publication timestamp. Four built-ins are registered in the inspected snapshot: a base PyTorch-plus-CUDA environment, an Ollama inference server, a ComfyUI image-generation interface, and an Ethen vLLM golden snapshot.
Your checklist entry mirrors the record shape. For each hosted candidate, demand: a pinned version, an immutability guarantee, a declared template or image, and a provenance label you can trace. If a provider cannot tell you exactly which version you deployed or whether that version can change, you cannot reproduce the deployment — and reproducibility is the hosted equivalent of Check 1's capability matrix. Note the disabled-record pattern too: availability is a per-recipe flag, not a property of the registry, which leads directly to the next two checks.
Check 6: How will you know the deployment is healthy?
Provisioning a GPU instance and serving a model are different achievements, and the recipe registry keeps them distinct through per-recipe health checks. Each record declares a protocol (TCP or HTTP), a port, and timing: a check interval plus a maximum retry count before the deployment is marked failed. HTTP checks add a path and an expected status range. These details differ per recipe in instructive ways.
The base environment checks TCP on port 22 — it verifies the machine is reachable, nothing more. The Ollama recipe checks HTTP on port 11434 at /api/tags and expects a 2xx status, which verifies the inference server is actually answering. ComfyUI checks HTTP on port 8188 at / with a wider 2xx–3xx acceptance range, matching a web interface that may redirect. The vLLM snapshot checks HTTP on port 8000 at /health with twelve-to-twenty-four retries at ten-to-fifteen-second intervals depending on the recipe — slower-starting servers get more patience. Expected ports and minimum disk sizes are declared alongside: SSH plus one service port in each record, and 50 GB for the base image versus 100–150 GB for the model-serving recipes.
The durable checklist has four lines. What does the health check prove — machine reachability or service readiness? Which port, path, and status range define success? How much startup patience do retries and intervals allow? And what disk floor does the recipe require before anything is downloaded? A TCP check on the SSH port is a legitimate answer for a base image and an insufficient one for an inference server; match the check to the claim. And as with the local path, nothing here establishes speed, cost, or output quality — a passing health check proves the service answers, not that it answers well.
Check 7: What does "available" actually mean?
The final check is about the word "available," which carries at least three distinct meanings that the registry keeps separate: the record exists, the record is enabled, and the backing artifact exists. The inspected snapshot demonstrates why the distinction matters. The vLLM golden-snapshot recipe is registered with a full specification — ports, health check, disk floor, recommended GPU type — but its enabled flag is false, with a comment noting golden snapshots are not yet available. The record is real; the deployment path is not currently offered. Conflating "registered" with "deployable" would misread the code.
Provenance strings add a second availability dimension. The three enabled recipes trace to provider templates, while the disabled one traces to an Ethen golden pipeline. Recommended GPU types appear only where the recipe author declared one. None of this describes fleet capacity, regional presence, pricing, or performance — those are live operational facts a static registry cannot carry, and the inspected code does not pretend otherwise.
Your checklist entry: for each hosted candidate, separate the recipe's existence from its enabled state, from the backing artifact's readiness, from live capacity. Ask which provenance each recipe traces to and whether the recommended hardware is a recorded suggestion or a measured requirement. This check is also where launch discipline matters most. Implementation evidence — a recipe record, a health check, an adapter — describes what the code models. It is not a release or availability announcement, and this guide makes no claim about which hosted paths are publicly offered. Verify current offering status against a primary source at decision time, or leave the question explicitly unresolved.
Putting it together: three walkthroughs
Three short scenarios show how the checks combine. Each starts from required operations and ends with a narrowed field — not a purchase recommendation, since cost and performance need measurements this guide does not supply.
First, a writing assistant that drafts offline. Required operations: status checks and streaming chat, nothing else. Every known local runtime record supports both, so Check 1 passes everywhere and the decision moves to Check 2: pick the endpoint dialect your client already speaks, native or OpenAI-compatible, and confirm the app boundary in Check 4 exposes exactly those two channels. No hosted recipe is needed, and install/delete support is irrelevant because the model set rarely changes.
Second, an experimentation bench that tries many open-weight models. Required operations: install, delete, detail inspection, and distinguishing running from installed models. Check 1 narrows the local field to runtimes whose APIs expose lifecycle operations — in the inspected matrix, only the Ollama record qualifies. Check 3 confirms lifecycle ownership lives in the runtime API rather than a separate GUI, and Check 4 adds confirmation dialogs and cancellation as requirements for safe high-churn use. If the local lifecycle story is insufficient, the hosted alternative must clear Checks 5 through 7 with a pinned, enabled recipe and a service-level health check.
Third, a shared team endpoint for one model behind an HTTP API. Local checks matter less here; the hosted checks dominate. Check 5 demands a pinned, immutable recipe with declared provenance. Check 6 demands an HTTP health check against the serving port and path — not merely TCP reachability — with retry patience matched to the server's startup time, plus a confirmed disk floor. Check 7 separates the recipe's registered state from its enabled state and the backing artifact's readiness. If any of those answers is missing or marked disabled, the path is not ready regardless of how complete the surrounding documentation looks.
Evidence and limitations
Every technical claim above traces to one of three inspected files. The capability matrix — provider kinds, per-operation flags, endpoint-dialect flags, conservative unknown-provider defaults, and the GUI-ownership reason strings — comes from the runtime capabilities module. The IPC boundary — localhost-only endpoints, model-name validation, main-process confirmation dialogs with their acknowledgement text, the request registry with cancellation, the single event channel, the missing-catalog limitation, and the fetch and chat timeouts — comes from the desktop IPC module. The recipe model — immutability, template-at-creation selection, health check specifications, expected ports, disk floors, enabled flags, provenance strings, and the four built-in records including the disabled golden snapshot — comes from the recipe registry. Product-surface separation (Desktop owns Local Models) is taken as a design target from the batch authority, and internal-link destinations were verified against the inspected route registry and marketing index rather than live HTTP.
The limitations are equally explicit. The code reflects the inspected snapshot and will evolve; capability flags, built-in recipes, and enabled states should be re-verified before acting. This guide makes no claims about inference speed, power consumption, privacy properties, or pricing on any path — those require fresh measurements scoped to specific hardware and workloads. Passing health checks prove reachability and declared readiness, not output quality. Registered recipes prove the code models a deployment, not that any hosted offering has launched. Where a fact could not be grounded in an inspected source, it was omitted rather than asserted.
The checklist, kept short
Choose the execution path by working these checks in order, and keep the answers with the decision so it stays reviewable: required operations listed per path; each operation verified against the runtime's capability record; endpoint dialect matched to every client; model lifecycle ownership assigned to exactly one surface; application boundary rules confirmed for validation, confirmation, cancellation, and timeouts; hosted recipes pinned, immutable, and provenance-labelled; health checks matched to the readiness claim with ports, paths, and patience declared; and enabled state separated from registered state and backing-artifact readiness. Paths that clear every applicable check are comparable on measurements you take next. Paths that fail an early check are out, whatever their other virtues.