Skip to content

EthenEthenEthen

Why Ethen Is Built Around Multiple Models Instead of One

Ethen is built around many AI models instead of one because no single model covers everything real work needs. Work spans different media — text, code, images, video, audio and speech — and the models that are best at each are different. Within any one kind of task, the strongest model is rarely also the fastest and cheapest. Some data may only be processed by certain models, or only on your own machine. Checking an AI's work is more trustworthy when the check does not come from the same model that did the work. And every model is eventually replaced, while the work people do with Ethen should outlast any one of them. A multi-model design lets Ethen fit each of those needs. To keep it from becoming chaos, Ethen relies on a shared body of sourced model knowledge, rules-first selection, and a clear record of which model did what.

Ethen is built around many AI models instead of one because no single model covers everything real work needs. Work spans different media — text, code, images, video, audio and speech — and the models that are best at each are different. Within any one kind of task, the strongest model is rarely also the fastest and cheapest. Some data may only be processed by certain models, or only on your own machine. Checking an AI's work is more trustworthy when the check does not come from the same model that did the work. And every model is eventually replaced, while the work people do with Ethen should outlast any one of them. A multi-model design lets Ethen fit each of those needs. To keep it from becoming chaos, Ethen relies on a shared body of sourced model knowledge, rules-first selection, and a clear record of which model did what.

Key takeaways

  • Different work needs different models. No single model is best across media, speed, cost and constraints.
  • Models are complementary. Large-scale evaluations find that different models succeed on different tasks.
  • Independent checking needs independence. A model tends to favor its own outputs, so judgments benefit from a different model family.
  • Work should outlive models. Projects, memory and evidence belong to Ethen and its users, not to any single model.
  • Choice must be explainable. Multi-model only works with sourced model facts, clear rules and a record of which model answered.

Why not just use the best model for everything?

Using the best model for everything sounds simple, but "best" changes depending on what you ask. The model that writes the most careful analysis may be slow and expensive for a quick classification. The model that generates the best images does not write code. The model that is excellent at long reasoning may not accept audio. And the model that is best today will be overtaken, repriced or retired.

There is also evidence that models genuinely complement each other. A large re-evaluation of model routing across hundreds of thousands of examples, many datasets and dozens of models confirmed strong complementarity among models — different models succeed on different items — even though it also found that many methods for exploiting that complementarity perform similarly and that several fail to reliably beat simple baselines (Li et al., 2026). Both halves of that finding matter: there is real value in having several models, and capturing that value takes discipline rather than cleverness. Earlier work showed that combining cheaper and more expensive models can reduce cost while preserving quality on the tasks studied (Chen et al., 2023).

What needs does a multi-model design cover?

A multi-model design covers five needs that one model cannot.

Table of five needs with why one model falls short and what multiple models allow.
Figure 1. Multiple models are not a hedge. They are how a product covers different media, constraints and kinds of checking.

Many media

Ethen spans conversation, code, research, images, video, audio and voice. The models that lead in each medium are specialists. Ethen Studio, for example, is the home for a large creative catalog organized by what each model does — text to image, image editing, text to video and so on — because creative work needs different models for different transformations. Even within one provider's catalog, many endpoints turn out to be variants of a smaller set of model families, which is why Ethen normalizes them before presenting choices; see From Provider Endpoints to Ethen Model Families.

Different trade-offs

Quality, speed and cost pull against each other. A product that serves everyday questions and hour-long analyses with the same model either overpays for the first or underserves the second. Multi-model lets Ethen use strong models where they matter and efficient ones where they are good enough. The right measure for that trade-off is the cost of good results rather than the price per token, a point we make in Making AI Model Choice Less Confusing.

Constraints

Some projects can only use models that meet specific data-handling or residency rules. Some work must stay on a person's own machine. A single-model product can only serve customers whose rules that one model happens to satisfy. Ethen's AI Gateway checks eligibility — capability, provider health, budget and policy — before choosing any model, and Ethen Desktop can run open models locally. Resilience across providers is a closely related topic, covered in Why Ethen Supports More Than One AI Provider.

Independent checking

When an AI system's work is checked by another AI model, independence matters. Research on using language models as judges has found that models can recognize and favor their own generations (Panickssery et al., 2024), alongside other biases such as sensitivity to the order in which answers are presented. Ethen Research Lab's methods paper Evaluating the Evaluators therefore recommends that judges from the same model family as the system being judged should not grade it, and that model judgments never override a failed deterministic check. It is a research synthesis. A single-model product cannot follow that advice; a multi-model one can.

Change over time

Every model is eventually replaced. When a product is built around one model, its prompts, its behavior and sometimes its users' work are entangled with that model. Ethen is designed so that projects, memory, evidence and the state of long-running work belong to Ethen and its users, and models are interchangeable components underneath. That separation is what makes it possible to adopt a better model — or decline one — on the evidence, which we discuss in Why Ethen Sometimes Won't Use the Newest Model.

Where do multiple models show up in Ethen?

Multiple models show up differently in each part of Ethen, on top of one shared body of model knowledge.

Five stacked bands: Chat, Studio, Code and Desktop, AI Gateway, Model Library and Model Intelligence.
Figure 2. Multiple models appear differently in each surface, on top of one shared body of model knowledge.

In Ethen Chat, you can pick a model or use an automatic option backed by Faros, Ethen's intelligence layer, which resolves to a default model for the conversation. In Ethen Studio, the full creative catalog is available, organized by what each model does. In Ethen Code and Desktop, cloud models sit alongside local models running on your own machine, so the choice can be made per task. The AI Gateway gives developers one key for many models, with eligibility checked before any selection. And the Ethen Model Library and Ethen Model Intelligence hold sourced facts about models so that choices can be explained.

How does Ethen keep many models from becoming chaos?

Ethen keeps many models from becoming chaos with three disciplines.

Sourced model knowledge. Model facts in Ethen have sources and owners, and Ethen Model Intelligence shows Unknown rather than guessing when provenance is missing. Choices are only as good as the facts behind them; we explain why in Why Ethen Is Investing in Model Intelligence.

Rules before learning. Ethen's Gateway selects by elimination: it removes models that cannot serve a request before ranking the rest, and it refuses rather than substituting something ineligible. Ethen Research Lab's survey Why Learned AI Model Routing Must Beat Good Rules argues that well-designed rules are the right baseline and that more sophisticated selection should be adopted only if it beats them. The survey reports no Ethen measurements; the Faros position paper frames the broader research question.

Stability within a task. Switching models in the middle of a chain of steps can break consistency and make results harder to explain. Model choices are kept stable within a task where possible, and switches happen at clear boundaries. And every result records which model produced it, so differences can be traced.

Where do open models fit?

Open models — models whose weights are published and can be run on your own hardware or a hosting provider of your choice — are an important part of a multi-model design, for reasons that have little to do with leaderboards. They can run locally, so data never leaves a person's machine. They can be deployed on dedicated hardware, so an organization controls the runtime, the version and the cost profile. Their behavior does not change behind a stable name unless the operator changes it. And they can be inspected and adapted in ways hosted models usually cannot.

The trade-off is that the strongest open models may trail the strongest hosted models on the hardest general tasks, and running them requires hardware and operations. That is exactly the kind of trade-off a multi-model design lets people make per task rather than once for everything. In Ethen, open models appear through local models on Desktop and through GPU deployment; we discuss the local case in Why Local AI Still Matters in a Cloud-First World.

What does multi-model mean for developers?

For developers building on Ethen, multi-model means designing applications that do not depend on one model's quirks. A few practices follow. Keep the state of your application — conversation history, files, task progress — in your own system rather than inside one provider's session format. Describe what a task needs (input types, tools, structured output, context length, data rules) explicitly, so eligibility can be checked rather than assumed. Test prompts and tool definitions on more than one model before relying on them. Record which model produced each result. And treat a model change as a change to be tested, not an automatic upgrade.

Ethen's AI Gateway is built to support those practices: one key across many models, requests bound to a project, eligibility checked before selection, and usage recorded per request. Whether a gateway is the right fit for a given application is covered in When an AI Gateway Helps.

Common misconceptions

"Multi-model means the system picks a random model." It should mean the opposite: a deliberate choice under explicit rules, recorded and explainable.

"Multi-model means always using the cheapest model." Cheaper models are used where they are good enough. Where quality matters, stronger models are used, and the comparison is the cost of good results, not of tokens.

"Multi-model is only about avoiding vendor lock-in." Lock-in is one reason. Different media, independent checking and work that outlives models matter at least as much.

How should a team adopt a multi-model approach?

A team does not need dozens of models to benefit from a multi-model approach. A small, deliberate set usually captures most of the value — large routing evaluations have also found that careful curation of a few models can beat larger collections (Li et al., 2026). A practical path has four steps.

Start with a default per kind of work. Pick one model for everyday writing and questions, one for heavier reasoning or code, and the specialist models your media work needs. Write down why each was chosen.

Add constraints explicitly. Mark which projects have data-handling, residency or local-only requirements, so ineligible models are excluded automatically rather than by memory.

Use a different model to check where judgment is involved. When an AI reviews an AI's work — a summary, an extraction, a draft — prefer a reviewer from a different model family, and keep deterministic checks such as tests or reconciliations wherever they exist.

Review the set on a schedule. New models, retirements and price changes are reasons to re-test against a saved set of your own tasks, not reasons to switch automatically.

That path keeps the benefits — fit, constraints, independent checking, longevity — while keeping the number of moving parts small enough to understand.

What does multi-model cost?

Multi-model has real costs. Models behave differently, so the same request can produce different styles and occasionally different conclusions. Every additional model is another configuration to test. Prompts and skills tuned for one model may perform worse on another. And people sometimes want one consistent voice rather than the best tool for each job. Ethen's answer is not to hide these costs but to manage them: let people pin a model when consistency matters, test changes on real work, and show which model produced each result.

An example

Illustrative example — hypothetical.

A product team prepares a launch. They use Ethen Chat to draft and refine the announcement, letting Ethen choose the model. They generate product images and a short video in Ethen Studio, where specialist image and video models are available. An engineer uses Ethen Code on Desktop to update the release notes, running a local model for a quick summary of internal changes that must not leave the laptop and a cloud model for a complex refactor. A research summary of competitor features is checked by a model from a different family than the one that wrote it, and two claims are flagged as unsupported. Five kinds of work, several models, one account — and a record of which model did what.

Frequently asked questions

Wouldn't one model give more consistent results? Sometimes, and for work that needs consistency you can pin a model. Across all kinds of work, one model cannot cover every medium, constraint and kind of check.

How does Ethen decide which model to use? By rules first: models that cannot serve a request or break its constraints are excluded before any ranking. In Chat, an automatic option resolves to a default model; you can also choose directly.

Is Faros a model? Faros is Ethen's flagship intelligence layer and first-party model identity. It is separate from the Ethen workspace and optional for model access.

Does using multiple models cost more? It can reduce cost by using efficient models where they are good enough, but it adds testing and management work. The useful comparison is the cost of good results.

References

  1. Li, H. et al. (2026). LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. arXiv:2601.07206. https://arxiv.org/abs/2601.07206
  2. Chen, L., Zaharia, M., Zou, J. (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176. https://arxiv.org/abs/2305.05176
  3. Panickssery, A. et al. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076. https://arxiv.org/abs/2404.13076
  4. Ethen Blog (2026). From Provider Endpoints to Ethen Model Families. https://upcube.ai/blog/from-provider-endpoints-to-ethen-model-families
  5. Ethen Research Lab (2026). Evaluating the Evaluators: Reward Integrity for AI Agents. Methods paper; research synthesis. https://upcube.ai/resources/research/reward-integrity
  6. Ethen Research Lab (2026). Why Learned AI Model Routing Must Beat Good Rules. External literature survey. https://upcube.ai/resources/research/learned-routing-vs-rules
  7. Ethen Research Lab (2026). Faros: Researching How Intelligence Should Choose Intelligence. Position paper. https://upcube.ai/resources/research/faros-research