What We Look for Before Adding a New Model to Ethen
Before adopting a new AI model, Ethen evaluates it against ten questions. Do we know exactly which model and version it is? Does it fill a gap — a medium, a task, a price or speed point, a local or open option — that Ethen cannot already fill well? Does it perform well on tasks like the ones people actually bring to Ethen, not only on public benchmarks? Is it reliable under real conditions? Will it change without notice? Do its data-handling terms and license allow the uses our users need? How does it behave with risky requests and untrusted content? Is its pricing clear and dated? Does it fit how Ethen calls models and tools? And where, if anywhere, should people see it? A model can be listed in our catalog with sourced facts long before it is qualified for a particular use, and qualified long before it is surfaced in a curated place like Ethen Chat. Each step requires its own evidence.
Before adopting a new AI model, Ethen evaluates it against ten questions. Do we know exactly which model and version it is? Does it fill a gap — a medium, a task, a price or speed point, a local or open option — that Ethen cannot already fill well? Does it perform well on tasks like the ones people actually bring to Ethen, not only on public benchmarks? Is it reliable under real conditions? Will it change without notice? Do its data-handling terms and license allow the uses our users need? How does it behave with risky requests and untrusted content? Is its pricing clear and dated? Does it fit how Ethen calls models and tools? And where, if anywhere, should people see it? A model can be listed in our catalog with sourced facts long before it is qualified for a particular use, and qualified long before it is surfaced in a curated place like Ethen Chat. Each step requires its own evidence.
Key takeaways
- Fill a gap, don't add a name. A new model should do something Ethen cannot already do well.
- Real tasks over leaderboards. Benchmarks are context; tests on representative work decide.
- Reliability and change policy matter as much as quality. A great model that changes without notice or fails under load is a liability.
- Terms are part of the model. Data handling, retention, residency options and licenses decide what users may do with it.
- Listed, qualified and surfaced are different. Appearing in a catalog is not the same as being ready for work.
Why does Ethen evaluate models before adding them?
Ethen evaluates models before adding them because adding a model is a promise. When a model appears in Ethen Chat's selector, in Ethen Studio's catalog, or as an eligible option in the AI Gateway, people reasonably assume it works for the purpose it is shown for, behaves predictably, and comes with terms they can live with. The pace of model releases makes it tempting to add every new model as soon as it appears. The cost of that temptation is a catalog full of options that silently fail, change behavior, or carry terms nobody read.
We have already written about one consequence of taking this seriously: in Ethen Studio, a model appearing in the catalog does not make it runnable; only a qualified route with evidence behind it does. See Building Durable Image and Video Jobs in Ethen Studio. The ten questions below generalize that principle.
1. Identity: exactly which model is this?
The first question is basic and often skipped: which model, from which publisher, in which family, at which version? Providers expose many endpoints that turn out to be variants of the same underlying model, and names change. Ethen normalizes provider endpoints into model families and records each fact with an owner and a source; see From Provider Endpoints to Ethen Model Families and Who Owns Each Model Fact in Ethen. If we cannot pin down what a model is, we cannot say anything reliable about how it behaves.
2. Gap: what does it do that we can't already do well?
A new model should fill a gap. Gaps come in a few kinds: a medium Ethen does not yet serve well; a task where existing models underperform; a speed or cost point that makes a workflow practical; an open or locally runnable option for people who need one; or eligibility under data rules that existing options cannot meet. A model that is slightly better on average than one we already have may still be worth adding — or may add cost and testing burden for little benefit. "It is new" is not a gap.
3. Quality on real tasks
Public benchmarks are useful context and poor deciders. They measure narrow skills on public tasks; their results are collected under varying methods; and they age as test items leak into training data, inflating scores without real capability gains (Zhang et al., 2024). Holistic evaluation efforts exist precisely because single numbers hide trade-offs across scenarios and metrics (Liang et al., 2022).
So the decisive test is how a model performs on tasks representative of the work people bring to Ethen, for the specific use we are considering — drafting, extraction, code changes, image editing, voice. Some of those tests use tasks we keep private, so that they remain a fair test of models that may have been trained on public data. When we cite a benchmark result publicly, it carries its source, date and method, or Ethen Model Intelligence shows Unknown; see Showing Unknowns in Ethen Model Intelligence.
4. Reliability under real conditions
A model that answers well in a demo can still fail in practice. We look at error rates and availability, latency and how much it varies, rate limits, how consistently it follows required formats such as structured output, and how reliably it calls tools. Reliability across repeated attempts matters as much as average quality: a model that succeeds on a task two times out of three is not ready for unattended work, even if its best output is excellent. Agent evaluation research has proposed measuring exactly this — the probability of succeeding on every one of several attempts — because averages hide inconsistency (Yao et al., 2024).
5. Change policy: will it change without notice?
Models change. Providers release new versions, retire old ones and sometimes alter behavior behind a stable name; a study of a widely used hosted model documented substantial behavior changes between versions released months apart (Chen et al., 2023). We look for version identifiers we can pin, clear deprecation timelines, and notice before changes. A model whose behavior can shift silently is harder to support, because every workflow built on it can break without warning. How we handle new versions is the subject of Why Ethen Sometimes Won't Use the Newest Model.
6. Terms and license
A model's terms are part of the model. For hosted models, that means how prompts and outputs are handled, whether and how long they are retained, whether they may be used for training, and which regions processing can happen in. For open-weight models, it means the license: whether commercial use is permitted, what restrictions apply, and what obligations come with redistribution or modification. These terms decide which projects a model is eligible for. A model can be excellent and still unsuitable for a customer whose data rules it cannot meet.
7. Safety behavior
We look at how a model handles risky requests, how it responds to instructions embedded in content it reads, and whether its content policies fit the uses we intend. This is not a certification; it is a check that a model's behavior is understood well enough to decide where it belongs. A model that is fine for drafting may need more guardrails before it is trusted to act on tools. Risk-management frameworks for AI recommend understanding a system's behavior in its context of use before deployment rather than assuming it from general reputation (NIST AI RMF 1.0).
8. Cost clarity
Pricing should be clear, published and dated, so that cost can be estimated before work runs and compared fairly across models. We record prices with their effective dates, because prices change, and we compare models on the cost of good results rather than price per token: a cheaper model that needs more retries or corrections may cost more per finished task.
9. Integration fit
A model has to fit how Ethen calls models: request and response conventions, streaming behavior, tool-calling formats, structured output support, context limits and error semantics. Differences here are not cosmetic. A model that formats tool calls differently or reports errors ambiguously can break workflows that depend on precise behavior, and those differences have to be handled before the model is offered for that use.
10. Placement: where should people see it?
The last question is where, if anywhere, people should encounter the model. Ethen has several kinds of place. Ethen Chat offers a deliberately small, curated set of choices plus an automatic option. Ethen Studio holds the full creative catalog. The AI Gateway can make a model available to developers without featuring it anywhere. And the public Ethen Model Library has pages only for model families with enough sourced information to say something useful. A model can be valuable in one place and wrong for another; putting it everywhere by default would make every surface noisier.
What happens after a model is added?
After a model is added, it is monitored. Reliability is watched, behavior is re-checked when versions change, prices are updated with their dates, and models that no longer meet the bar — because they degraded, were superseded, or changed their terms — are moved out of curated places or retired. Ethen Research Lab has proposed a more systematic version of this practice, Model Change Assurance: before a model change reaches live work, replay a sample of real past tasks under the old and new configurations and compare them task by task. It is a research note describing an architecture proposal; the protocol to validate it has not been run.
An example
Illustrative example — hypothetical, not a specific model.
A new speech model is released with impressive demo audio. Identity: it is a new family from an established publisher, with a pinned version. Gap: it offers natural-sounding narration in several languages Ethen serves less well. Quality: on a set of representative narration scripts, including some kept private, it sounds strong in four languages and mispronounces names in two. Reliability: latency is acceptable but varies widely at peak times. Change policy: versions are pinned and deprecations announced. Terms: audio inputs may be retained for a period; voice cloning requires consent documentation. Safety: it will generate a voice from a short sample, so consent checks are essential. Cost: clear, dated per-character pricing. Integration: streaming works with minor adaptation. Placement: qualified for narration in four languages in Ethen Studio, with voice cloning limited to consented voices; not featured in Chat's curated list; re-tested in the two weaker languages on the next version.
How can your team apply this?
Most teams do not run a model catalog, but almost every team adopts new models. The same questions work at smaller scale, and a lighter version takes an afternoon rather than a quarter.
Write down the gap. One sentence on what the new model should do better than what you use now. If you cannot write it, you probably do not need the model.
Test on twenty to fifty real examples. Use your own inputs, including awkward ones, with pass criteria written before you run anything. Run each example more than once if outputs vary.
Read the terms before the benchmarks. Data retention, training use, processing regions and license restrictions can disqualify a model faster than any quality test.
Check the change policy. Can you pin a version? Will you be told before it changes or is retired?
Decide where it goes. Not every good model should replace your default. Some belong behind a specific workflow, some in an API-only role, and some only in a trial.
Set a re-check date. Put a reminder to re-run your test set when the model updates or prices change.
The result is a short record that explains why a model was adopted and how to tell if that decision has stopped being right — which is most of the value of a formal evaluation, at a fraction of the effort.
Tradeoffs
Careful evaluation means new models sometimes reach Ethen later than they reach other places. Some people want every model immediately, and for them a broad catalog with less curation is a reasonable choice elsewhere. We think the cost of a slower, more deliberate catalog is worth it for people who need models to work, behave predictably and come with terms they can rely on. We also try to keep the gap small by listing models with their sourced facts early, while qualification for specific uses follows.
Frequently asked questions
Does Ethen add every new model? No. A model is added when it fills a gap and passes checks for quality on real tasks, reliability, change policy, terms, safety, cost and integration. It can be listed with sourced facts before it is qualified for specific uses.
Why isn't a model with great benchmark scores available in Ethen Chat? Chat offers a small, curated set of choices. A model may be available elsewhere — in Studio or through the Gateway — without being featured in Chat.
Does Ethen test models on its own tasks? Yes. Representative tasks, including some kept private so they remain a fair test, matter more than public benchmarks for deciding where a model belongs.
What happens when a model changes or degrades? It is re-checked. Models that no longer meet the bar move out of curated places or are retired.
Related reading
- What Makes an AI Model Useful Beyond Benchmarks
- How to Read an AI Model Comparison
- Model Intelligence and Routing
- Ethen Model Intelligence
References
- Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332
- Liang, P. et al. (2022). Holistic Evaluation of Language Models. arXiv:2211.09110. https://arxiv.org/abs/2211.09110
- Yao, S. et al. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045. https://arxiv.org/abs/2406.12045
- Chen, L., Zaharia, M., Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv:2307.09009. https://arxiv.org/abs/2307.09009
- National Institute of Standards and Technology (2023). AI Risk Management Framework (AI RMF 1.0), NIST AI 100-1. https://doi.org/10.6028/NIST.AI.100-1
- Ethen Blog (2026). Building Durable Image and Video Jobs in Ethen Studio. https://upcube.ai/blog/building-durable-image-and-video-jobs-in-ethen-studio
- Ethen Blog (2026). Showing Unknowns in Ethen Model Intelligence. https://upcube.ai/blog/showing-unknowns-in-ethen-model-intelligence
- Ethen Research Lab (2026). Model Change Assurance: Testing AI Upgrades Before They Reach Real Work. Research note; architecture proposal. https://upcube.ai/resources/research/model-change-assurance