Versioning GPU Deployment Recipes in Ethen
Ethen GPU deployment recipes pin each GPU launch to a named, versioned template with explicit ports, disk, and health checks — and keep newer ideas disabled until the backing image exists.
Ethen GPU deployment recipes pin each GPU launch to a named, versioned template with explicit ports, disk, and health checks — and keep newer ideas disabled until the backing image exists.
GPU deployments drift. A working inference server stops working because the base image changed, a port moved, or a dependency updated silently underneath it. Ethen's answer in the inspected infrastructure code is to make the deployment description a versioned, immutable record called a recipe, select it once at instance creation time, and verify the result with a per-recipe health check. Nothing in this design runs ad-hoc remote commands at launch.
This article traces that mechanism as implemented in two files: a recipe registry that defines and stores recipes, and a provider adapter that talks to one compute integration. It also explains the sharpest edge in the current data: the Ethen vLLM golden snapshot recipe exists in code but is explicitly disabled.
What a recipe actually contains
A recipe is a struct, not a script. Each record carries a stable identifier such as ollama, comfy-ui, or ethen-vllm-v3, a human-readable name, a semantic version string, and a description of what the recipe provides. The operational fields are what make it deployable:
templateName: the backing provider template name or golden snapshot name.kind: eitherprovider_templateorgolden_snapshot.expectedPorts: the ports that should be open after deployment.healthCheck: a protocol, port, and polling policy described below.enabled: whether the recipe is available for use.provenance: a source label such asthunder-templates:ollamaorethen-golden-pipeline:v3.minimumDiskSizeGb: the minimum disk required.recommendedGpuType: an optional hint such asL40-48GB.publishedAt: a timestamp set once the recipe version is published.
The header comment in the registry states the intent directly: recipes are versioned and immutable, backed by provider templates or golden snapshots. That framing matters because it narrows what a recipe promises. It does not promise that a model will load, that throughput will meet a target, or that any particular GPU fleet is reserved. It promises that a specific id plus version always resolves to the same template name, ports, health check, disk floor, and provenance.
The RecipeKind union has only two members. There is no third path for custom inline setup, no shell hook, and no per-tenant override in the type. If a workload needs something different, the registry's model says to publish a new recipe version with a new template name — not to mutate the existing record.
Selection happens once, at creation time
The registry's second design rule is about timing. Recipes are selected as the template field at instance creation time, and no per-launch remote command execution occurs.
In practice that means the choice of ollama@1.0.0 versus comfy-ui@1.0.0 is a single value passed when the instance is requested. There is no second step where Ethen SSHes in and installs packages, pulls a different server, or rewrites configuration based on who is launching. The template backing the recipe is expected to already contain what the workload needs.
This is a meaningful constraint for operators. It removes an entire class of launch-time variance — failed downloads, transient package mirrors, interactive prompts — but it also moves responsibility earlier. Whoever builds or names the template must get it right, because the launch path will not fix it up. The provenance field records where that template came from, so a reader can distinguish a native provider template from an Ethen-built golden snapshot without guessing from the recipe name alone.
An example helps. Suppose a caller wants the Ollama server. It resolves recipe id ollama, version 1.0.0, reads templateName: "ollama", and passes that template name into the provider's instance-create request. The registry does not itself call the provider; it only answers the question "what does this version mean?" The adapter described later is the part that speaks to the provider API.
Immutability is enforced, not suggested
The registry stores recipes in a module-level Map keyed by id@version. Registration goes through registerRecipe, which throws if the same id plus version is already registered:
Recipes are immutable once published.
The stored object is also frozen with Object.freeze after a shallow copy, so later code cannot mutate the registered record in place. The publishedAt field is documented the same way: immutable once set.
Retrieval follows the same keying:
getRecipe(id, version)returns one exact version or null.getLatestRecipe(id)scans keys with the prefixid@and keeps the entry whoseversionstring compares greater.listRecipes()returns the latest version for each distinct id.listAvailableRecipes()filters that list toenabledrecipes.hasRecipe(id, version?)checks for a specific version or for any version of an id.
Two implementation details deserve attention because they bound what callers should assume.
First, "latest" is a lexicographic string comparison (recipe.version > latest.version), not a parsed semantic-version comparison. For the current built-in data this distinction does not change the answer — each id has one version — but a future set such as 1.0.10 versus 1.0.9 would not necessarily order numerically. Treat the helper as "greatest version string," not as a guarantee of semver-aware resolution.
Second, listRecipes returns one row per id. It is a catalog view, not an audit log. Older versions remain in the map and remain retrievable by getRecipe, but they do not appear in the list output. That matches the versioning story — old versions keep working for existing references while new launches default to the latest — but anyone building a history UI needs to know the list helper alone will not show prior versions.
The four built-in recipes
registerBuiltinRecipes seeds four records, all stamped 2026-07-28T00:00:00.000Z. Three are enabled provider templates; one is a disabled golden snapshot.
Base 1.0.0 is the minimal environment: PyTorch plus CUDA support for custom workloads. It expects ports 22 and 8080, requires at least 50 GB of disk, and checks TCP connectivity to port 22 every 10 seconds with up to 12 retries. Its provenance is thunder-templates:base.
Ollama 1.0.0 is the Ollama inference server for running open-source language models behind one API call. It expects ports 22 and 11434, requires at least 100 GB, and checks HTTP GET /api/tags on port 11434 for a 2xx response every 10 seconds with up to 18 retries. Provenance is thunder-templates:ollama.
ComfyUI 1.0.0 is the node-based image generation interface. It expects ports 22 and 8188, requires at least 100 GB, recommends an L40-48GB GPU type, and checks HTTP GET / on port 8188 for a 200–399 response every 15 seconds with up to 20 retries. Provenance is thunder-templates:comfy-ui.
Ethen vLLM 3.0.0 is the outlier. It describes an Ethen golden snapshot with a vLLM inference server, backed by template name ethen-vllm-v3, expecting ports 22 and 8000, checking HTTP GET /health on port 8000 for 2xx every 15 seconds with up to 24 retries, requiring 150 GB and recommending L40-48GB. Its provenance is ethen-golden-pipeline:v3. And its enabled flag is false, with the inline comment "Golden snapshots not yet available."
That flag is load-bearing. listAvailableRecipes filters it out, so any caller that lists what is launchable will see three recipes, not four. The record is still registered and still retrievable by exact id and version, which preserves the version identifier for future work, but it is not offered as a current choice. Do not describe it as available, supported, or launched.
The disk and GPU hints also deserve a literal reading. minimumDiskSizeGb is a floor per recipe — 50, 100, 100, and 150 GB respectively — not a measured usage report. recommendedGpuType appears only on ComfyUI and Ethen vLLM, both naming L40-48GB. The field is optional, and its name says recommended, not required. The registry records the suggestion; it does not enforce scheduling or prove that such capacity is on hand.
Health checks are the deployment contract
Every recipe carries a HealthCheckSpec. The protocol is either tcp or http. Both variants share a port, a poll interval in milliseconds, and a maximum retry count. HTTP adds a path and an expected status range.
The built-in values show how the contract tightens as the service surface gets more specific:
| Recipe | Protocol | Target | Success | Interval | Retries |
|---|---|---|---|---|---|
| base 1.0.0 | tcp | port 22 | connection accepted | 10,000 ms | 12 |
| ollama 1.0.0 | http | port 11434 /api/tags | 200–299 | 10,000 ms | 18 |
| comfy-ui 1.0.0 | http | port 8188 / | 200–399 | 15,000 ms | 20 |
| ethen-vllm 3.0.0 | http | port 8000 /health | 200–299 | 15,000 ms | 24 |
The base recipe checks only SSH-level reachability. That confirms the machine answered, not that any inference server started — which is appropriate for a generic PyTorch and CUDA environment with no single service to probe. Ollama checks a real API route (/api/tags) and requires a 2xx, which confirms the server is up and serving its tag list. ComfyUI checks / and accepts 2xx through 3xx, a wider range that tolerates redirects on a browser-facing UI. The disabled vLLM recipe checks a dedicated /health endpoint with the longest retry budget, 24 attempts at 15-second intervals, consistent with a heavier server that may take minutes to load weights.
To translate retries into wall-clock patience: base allows about 2 minutes of polling, Ollama about 3 minutes, ComfyUI about 5 minutes, and vLLM about 6 minutes. Those are ceilings implied by multiplying interval by retries, not measured startup times. The registry does not record how long any real deployment took.
What the health check does not do is equally important. It does not authenticate, does not run inference, and does not verify model identity or output quality. A 200 from /api/tags proves the Ollama server responded; it does not prove which models are pulled. A 200 from ComfyUI's / proves the interface answered; it does not prove any workflow ran. Readers evaluating GPU hosting should treat these checks as liveness gates — necessary, useful, and insufficient for correctness.
Provider templates versus golden snapshots
The kind field separates two supply chains. provider_template means a native template supplied through the provider's template catalog. All three enabled recipes use this kind, and their provenance strings share the thunder-templates: prefix followed by the template name. golden_snapshot means an Ethen-built snapshot produced by an Ethen pipeline. Only the disabled vLLM recipe uses this kind, with provenance ethen-golden-pipeline:v3.
That naming is provenance, not marketing. It tells the operator where to look when the template misbehaves: provider-side catalog versus Ethen-side snapshot pipeline. It does not assert ownership of hardware, a partnership with the provider, or any service-level agreement. A provenance string is a label in a struct; fleet and commercial relationships would need contracts, capacity data, and operational evidence that these two files do not contain.
The golden-snapshot path is best read as a design target in this snapshot of the code. The type supports it, the vLLM record sketches what it would look like — template name, health check, disk floor, GPU hint — and the enabled flag plus the "not yet available" comment say it is not ready to use. Versioning makes that posture cheap: ethen-vllm@3.0.0 reserves the identifier and documents the intended contract while staying out of the available list. When and whether a golden pipeline produces a bootable, verified snapshot is unverified from these sources alone.
What the Thunder adapter does — and does not prove
The second inspected file is the Thunder provider adapter, marked R01. Its header states the layering rule: it implements the neutral ComputeProvider interface by delegating to the Thunder integration layer, and it is the only module that imports from that layer. That single-import boundary is the point. Everything above the adapter speaks provider-neutral types; only the adapter knows Thunder request shapes, response normalizers, and snapshot helpers.
The adapter's surface covers the instance and snapshot lifecycle:
- Instances:
createInstance,listInstances,getInstance,modifyInstance,deleteInstance. Note thatgetInstanceis implemented by listing all instances and finding a matching id, not by a dedicated get-by-id API call. That works but inherits whatever consistency and scale limits the list operation has. - Snapshots:
createSnapshot,getSnapshot,findSnapshotsByName,deleteSnapshot,pollSnapshotReady,confirmSnapshotDeletion.getSnapshottolerates provider list payloads that surface eitheridorsnapshotId, andpollSnapshotReadymaps a providerreadystatus to neutralCOMPLETED, with other outcomes surfaced as the raw status orTIMEOUT. - Ports:
exposePortscallsmodifyInstancewith the requested ports inprototypingmode;removePortsclears exposure by passing an empty list. The recipe'sexpectedPortsand this port-exposure call are related but distinct: the recipe declares what should be open, while the adapter performs the provider operation that opens them. - Deletion confirmation:
confirmDeletionandconfirmSnapshotDeletionpoll until the provider confirms removal, converting an optional timeout into a retry count.
Two safety details are worth naming. First, deleteSnapshot carries an explicit guard comment: snapshot deletion must call the provider snapshot-delete operation, never the instance-delete endpoint, because routing a snapshot id to instance deletion could destroy a live instance. Second, error classification is centralized in classifyError plus a code mapper. Thunder error codes translate to neutral codes — invalid credentials to authentication, provider rejection and schema validation to validation, provider unavailability to retryable, rate limiting to rate-limited — while message sniffing maps timeouts to indeterminate, 5xx and network failures to retryable, and anything else to non-retryable.
None of this proves a partnership, an owned fleet, or production readiness. An adapter is integration code: it shows Ethen knows how to shape requests for one provider's API and how to normalize responses. It does not show a signed agreement, reserved capacity, measured uptime, or any live deployment. The file contains no fleet inventory, no credentials, no endpoint URLs, and no provisioning logs. Readers should treat it as evidence of a software boundary, not as evidence of infrastructure ownership.
Why the vLLM recipe stays disabled
It is tempting to read ethen-vllm@3.0.0 as the flagship — highest version number, purpose-built description, generous retry budget — and conclude it is the recommended path. The code says the opposite. enabled: false removes it from every availability listing, and the accompanying comment gives the reason: golden snapshots are not yet available.
There are good engineering reasons to keep a record like this registered but disabled. The id and version are reserved, so future work cannot accidentally reuse ethen-vllm@3.0.0 for a different contract. The health check, ports, disk floor, and GPU hint document the intended shape, so the snapshot pipeline has a concrete target to hit. And because registration is immutable, publishing the disabled record now does not let anyone quietly change what version 3.0.0 means later; enabling it would require either flipping the flag in a new registration — which the immutability rule forbids for the same id plus version — or publishing a new version. Careful readers will note the tension: the current code as seeded cannot enable 3.0.0 in place without violating its own rule, which suggests the real enablement path is a future version such as 3.0.1 or 4.0.0 once the snapshot exists and is verified. That inference is structural, not stated, and should be labeled as such.
The practical consequence is simple. Anyone describing currently selectable Ethen GPU deployment recipes from this source must list three: base, Ollama, and ComfyUI. The vLLM golden snapshot is a documented future, not a current option. Launch-oriented claims about GPU compute — supported model deployments, availability, pricing, or performance — belong to a separate release gate and cannot be grounded in these two files.
Evidence and limitations
What the inspected sources establish:
- Recipes are versioned by
id@version, frozen on registration, and retrievable exactly or as latest-per-id. - Selection is a single template value at creation time with no per-launch remote execution.
- Each recipe declares ports, a health check with protocol, path, status range, interval, and retries, plus disk floor, optional GPU hint, provenance, and enabled flag.
- Three provider-template recipes are enabled; one golden-snapshot recipe is disabled with golden snapshots marked not yet available.
- A single adapter module isolates Thunder-specific calls behind a provider-neutral interface, with explicit snapshot-delete safety and a documented error-code mapping.
What they do not establish:
- No partnership, ownership, capacity, uptime, pricing, or geographic availability for any GPU fleet.
- No verified end-to-end provisioning run, cleanup run, or measured startup time.
- No enforcement of
minimumDiskSizeGborrecommendedGpuTypeat scheduling time — the registry records them; scheduling behavior is out of scope for these files. - No semver-aware ordering for "latest" beyond string comparison.
- No dedicated get-instance API;
getInstancescans the list response. - No model-level verification from health checks; liveness is not correctness.
The target-app architecture also stays out of this story's scope. Chat, Studio, Research, Designer, Founder, Computer, Desktop, and Code have separate ownership boundaries documented elsewhere, but nothing in the recipe registry or the Thunder adapter assigns GPU recipes to any of those surfaces. Do not infer which product uses which recipe from these files.
Choosing among the available recipes
For the three enabled recipes, the choice reduces to matching the workload to the declared contract.
Start with the service shape. If the goal is a custom PyTorch and CUDA environment without a fixed server, base 1.0.0's TCP check on port 22 and 50 GB floor reflect that generality. If the goal is an OpenAI-compatible-style language-model server with a tag API, Ollama 1.0.0's HTTP check on port 11434 is the tighter contract. If the goal is a visual generation workflow in the browser, ComfyUI 1.0.0's HTTP check on port 8188 and its L40-48GB hint point that way.
Then check the operational fit: expected ports against firewall and exposure policy, disk floor against quota, and health-check patience against startup expectations. A caller that gives up after 60 seconds will wrongly mark healthy ComfyUI or Ollama launches as failed, because the recipes budget two to five minutes of polling. Conversely, a caller that treats a passing TCP check as proof the application is ready will proceed too early on the base recipe, where only machine reachability is verified.
Finally, treat the disabled vLLM recipe as a signal to wait, not to work around. Its presence shows where Ethen intends to go — a purpose-built vLLM snapshot with a dedicated health endpoint — but its flag and comment show the snapshot pipeline has not delivered. Building automation that references ethen-vllm@3.0.0 today would pin to an explicitly unavailable contract.
What versioning buys
Immutable, versioned recipes trade flexibility at launch for certainty about what launched. The template name cannot drift under a pinned version, the health check cannot silently widen, and the disk floor cannot quietly drop. When something changes, it gets a new version, and old references keep their meaning. The Thunder adapter complements that certainty by keeping provider quirks — dual id fields, status vocabularies, error codes, port modes — behind one translatable boundary.
That is a modest but real foundation for GPU operations. It does not by itself deliver hosted models, guarantee capacity, or certify any deployment as production-ready. It does give every future claim about Ethen GPU deployment recipes a stable identifier to point at — and, for now, it points the vLLM story at a version that is documented, reserved, and deliberately switched off.