Skip to content

EthenEthenEthen

Why Ethen Is Investing in Model Intelligence

Ethen invests in model intelligence because every decision about which AI model to use — made by a person, by Ethen's AI Gateway, or by Ethen's automatic model choice — is only as good as the facts behind it. Model intelligence is the knowledge needed to make that decision: what a model can do, what it costs, how it performs on which kinds of task, and where each of those facts came from. The model landscape changes too quickly, and public comparisons hide too much, for that knowledge to be assembled ad hoc. So Ethen builds it deliberately: every fact with a source and an owner, Unknown shown instead of guesses, eligibility decided before preference, and a long-term aim of connecting model choices to whether the resulting work actually succeeded.

Ethen invests in model intelligence because every decision about which AI model to use — made by a person, by Ethen's AI Gateway, or by Ethen's automatic model choice — is only as good as the facts behind it. Model intelligence is the knowledge needed to make that decision: what a model can do, what it costs, how it performs on which kinds of task, and where each of those facts came from. The model landscape changes too quickly, and public comparisons hide too much, for that knowledge to be assembled ad hoc. So Ethen builds it deliberately: every fact with a source and an owner, Unknown shown instead of guesses, eligibility decided before preference, and a long-term aim of connecting model choices to whether the resulting work actually succeeded.

Key takeaways

  • Model choice is a knowledge problem before it is a ranking problem. Rankings built on weak facts are confident and wrong.
  • Provenance is non-negotiable. Ethen Model Intelligence shows Unknown when a value lacks a source, a retrieval date, a methodology or adequate confidence.
  • Each fact has an owner. When two systems disagree about a model, there is a defined answer about which one is responsible.
  • Eligibility comes before preference. The Gateway removes models that cannot serve a request before ranking the rest.
  • Knowledge about models outlasts any one model. Specific models churn; the discipline of knowing what is true about them does not.

What is model intelligence?

Model intelligence is the structured, sourced knowledge needed to choose an AI model for a task. It covers identity (which model this is, from which publisher, in which family), capabilities (input and output types, tool support, context limits), economics (prices and their effective dates), performance evidence (benchmark results with their methodology), availability, and known limitations. Model routing is the decision that uses this knowledge: which model, provider and configuration should handle a particular request.

At Ethen, model intelligence is a product and a foundation at the same time. Ethen Model Intelligence is where people compare models by benchmarks, pricing, speed and provider, with sources shown. The Ethen Model Library is the structured catalog of models and providers Ethen can use, and each model page is designed as a routing point into the rest of Ethen — try it in Chat, use it in Studio, run it locally, deploy it, or call it through the Gateway. Underneath both, the same facts feed the AI Gateway and Faros, Ethen's intelligence layer.

Why does choosing a model need dedicated investment?

Choosing a model needs dedicated investment because four things make it harder than it looks.

The landscape moves weekly. New models, new versions of existing models, price changes and deprecations arrive constantly. A comparison that was right in spring can be wrong by summer, and a decision made without a date attached cannot be re-checked.

Public comparisons answer narrower questions than their headlines. A leaderboard built on one benchmark speaks only to what that benchmark measures. Scores collected from different sources may use different methods, prompts or settings. Holistic evaluation efforts in the research community exist precisely because single numbers hide trade-offs across scenarios and metrics (Liang et al., 2022). And public benchmarks age: when test items leak into training data, scores rise without capability rising with them, as a carefully matched replacement for one widely used math benchmark showed (Zhang et al., 2024).

Different jobs need different models. The best model for a long reasoning task may be the wrong choice for a quick classification, a vision task, or a workload with strict data-residency requirements. "Best model" is almost always a question with a missing clause: best for what, under which constraints, at what cost?

Many decisions are automated. When Ethen chooses a model on someone's behalf — through the automatic option in Chat or through the Gateway — the quality of that choice depends entirely on the facts available to the system. A person can sanity-check a strange recommendation; an automated router will act on whatever it is given.

How does Ethen build model knowledge?

Ethen builds model knowledge around four principles, each described in a public engineering post.

1. Identity first. Before facts can be compared, models have to be identified correctly. Provider catalogs expose many endpoints that turn out to be the same underlying model family, and a catalog that treats each endpoint as a separate model will mislead anyone comparing them. Our catalog work normalizes endpoints into families and decides separately which families have enough information to deserve a public page. See From Provider Endpoints to Ethen Model Families.

2. Facts need provenance — or they are Unknown. In Ethen Model Intelligence, a benchmark value is shown only when it has a source URL, a valid retrieval date, a methodology statement and adequate confidence. If any of those is missing, the display shows Unknown and keeps the reason attached. It does not interpolate, carry forward an old score, or borrow a number from a similar benchmark. See Showing Unknowns in Ethen Model Intelligence.

Table of five required inputs for showing a benchmark value — value, source URL, retrieval date, methodology, adequate confidence — each shown as Unknown if missing.
Figure 1. Ethen Model Intelligence shows Unknown rather than a number it cannot support.

3. Every fact has an owner. A model page looks like one record, but its facts come from different places with different update rhythms: catalog identity changes slowly, runtime health changes by the minute, routing policy is decided at request time, and display formatting depends on the page. Ethen assigns each named model fact to exactly one owning system, so that when two systems disagree, it is clear which one is supposed to answer. See Who Owns Each Model Fact in Ethen.

4. Eligibility before preference. When the Gateway serves a request, it first removes every model that lacks a required capability, is unhealthy, or would exceed the budget, and only then ranks what is left. If nothing survives, it does not pick a fallback silently; it refuses, and the decision record says why. See How Ethen Gateway Chooses an Eligible Model.

Five stacked bands: Identity, Facts with provenance, Eligibility, Choice, Outcome.
Figure 2. Model intelligence is the lower half of every model decision. A ranking at the top cannot be better than the facts at the bottom.

An example: what good model knowledge changes

Illustrative example — a hypothetical decision, not a measured result.

A team wants a model to extract line items from scanned supplier invoices. A popular leaderboard ranks a large general model first. Without model intelligence, the team picks it and moves on.

With model intelligence, the decision looks different. Identity resolves the "first place" model to a family with several endpoints, only some of which accept image input. Capability facts show which models in the shortlist accept scanned documents and return structured output; two strong models drop out immediately. The leaderboard score turns out to come from a reasoning benchmark that says nothing about document extraction, so its cell for this task is effectively Unknown. Prices carry effective dates, revealing that one candidate's price changed last month. Eligibility under the team's policy removes a model whose provider cannot meet their data-residency requirement.

What remains is a short list of two or three models that can actually do the job under the team's constraints. The team runs a small test on fifty of its own invoices, checks the extracted totals against the source documents, and compares the cost of verified correct extractions rather than the price per token. The winner may or may not be the leaderboard's first place. Either way, the team knows why they chose it, and can re-check the decision when a new model or price appears.

Why not just publish a leaderboard?

Because a leaderboard is a conclusion, and model intelligence is the evidence a conclusion should rest on. Leaderboards are useful when they are honest about scope: which benchmark, which method, which date, which models were actually measured by whom. They become misleading when they compress all of that into one ordering and present it as "the best model".

Ethen Model Intelligence deliberately does less than a typical leaderboard in some ways. It does not publish an original Ethen index score. It shows empty cells where provenance is incomplete, even when a number is available somewhere. And it began narrow on purpose: our own guide to reading comparisons notes that the current evaluation registry represents a single benchmark with full identity, methodology and comparability evidence, which means its supported comparisons are correspondingly narrow. That is a scope statement, not an apology. We would rather expand a small set of trustworthy comparisons than publish a broad set we cannot defend. Our reader's guide How to Read an AI Model Comparison explains how to apply the same skepticism to any comparison, including ours.

Why does model knowledge last longer than models?

Specific models come and go. The knowledge of how to evaluate them, and the record of how they performed on real work, compounds. Ethen Research Lab has argued in its position paper Why Better Foundation Models May Make Evaluation More Valuable, Not Less that as models become more capable and cheaper, organizations switch models more often and give them longer tasks — which raises the value of knowing, quickly and accurately, which workflows a change will improve and which it will break. That paper is a research synthesis and argument; it reports no measured results.

The practical consequence for Ethen is a direction rather than a shipped feature: model intelligence should increasingly connect to outcomes. Benchmarks tell you how a model did on someone else's tasks. The most useful evidence is how a model, in a given configuration, performs on the kind of work you actually do — whether the work succeeded, what it cost, how often it needed correcting. Ethen Research Lab's note on Cost Per Verified Outcome proposes measuring cost per independently verified success rather than per token, because a cheaper model that needs more retries can make finished work more expensive. That too is a research proposal, and Ethen's own cost per verified outcome has not been measured.

How does this connect to Faros?

Faros is Ethen's intelligence layer and first-party model identity. It is separate from the Ethen workspace, and people can choose models directly or use the Gateway without it. In Chat, an automatic option resolves to a default model for the conversation. Ethen Research Lab's position paper Faros: Researching How Intelligence Should Choose Intelligence frames the long-term research question more broadly than model choice: given a task, its constraints and a budget, which complete configuration — model, context, tools, verification and recovery — is most likely to produce a verified result? The paper argues that strong rules are the right starting point and that any learned decision-making should have to beat them. It is a position paper with no Ethen routing measurements.

Whatever form automatic choice takes, it inherits the quality of the facts beneath it. That is the most direct reason Ethen invests in model intelligence: it is the part of every model decision that has to be right first.

What does this mean for people choosing models?

For someone choosing a model today, Ethen's investment translates into a few practical habits we try to support and encourage:

  • Start from the job. Write down what the task requires — modality, tools, context length, latency, data constraints — before looking at any ranking.
  • Filter before you rank. Remove models that cannot do the job or cannot meet your constraints; then compare the rest.
  • Read the provenance. Check the source, date and method behind any number you rely on. Treat a missing value as missing, not as average.
  • Prefer evidence from your own work. A small test on your own tasks often tells you more than a large public benchmark.
  • Re-check when things change. A new model version or price is a reason to re-evaluate, not an automatic upgrade.

We turn these habits into a practical method in Making AI Model Choice Less Confusing, and look at what benchmarks miss in What Makes an AI Model Useful Beyond Benchmarks.

Limitations

Model intelligence has real limits, and some of them are ours. Provider statements about capabilities and prices are not always independently verifiable, and they change. Benchmark coverage in Ethen Model Intelligence is deliberately narrow today, so many comparisons people want are not yet supported with full provenance. Showing Unknown is honest but can be frustrating when a decision has to be made anyway. And the connection between model choice and verified outcomes on real work is direction and research, not a shipped capability.

Frequently asked questions

What is Ethen Model Intelligence? It is Ethen's product for comparing AI models by benchmarks, pricing, speed and provider, with sources shown. It shows Unknown when a fact lacks adequate provenance.

Does Ethen publish its own model rankings? Ethen Model Intelligence does not publish an original Ethen index. It presents sourced evidence and its gaps, so readers can judge comparisons themselves.

How does Ethen choose a model automatically? The Gateway first removes models that cannot serve a request, then ranks eligible ones under policy. In Chat, an automatic option resolves to a default model. The Faros research program studies how such decisions should be made; its learned-routing work has not been run.

Why does a model page show Unknown for a benchmark? Because one of the required inputs — source, retrieval date, methodology or adequate confidence — is missing. A number without those would imply a comparison the evidence cannot support.

References

  1. Liang, P. et al. (2022). Holistic Evaluation of Language Models. arXiv:2211.09110. https://arxiv.org/abs/2211.09110
  2. Zhang, H. et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332. https://arxiv.org/abs/2405.00332
  3. Ethen Blog (2026). Showing Unknowns in Ethen Model Intelligence. https://upcube.ai/blog/showing-unknowns-in-ethen-model-intelligence
  4. Ethen Blog (2026). Who Owns Each Model Fact in Ethen. https://upcube.ai/blog/who-owns-each-model-fact-in-ethen
  5. Ethen Blog (2026). How Ethen Gateway Chooses an Eligible Model. https://upcube.ai/blog/how-ethen-gateway-chooses-an-eligible-model
  6. Ethen Blog (2026). From Provider Endpoints to Ethen Model Families. https://upcube.ai/blog/from-provider-endpoints-to-ethen-model-families
  7. Ethen Blog (2026). How to Read an AI Model Comparison. https://upcube.ai/blog/how-to-read-an-ai-model-comparison
  8. Ethen Research Lab (2026). Why Better Foundation Models May Make Evaluation More Valuable, Not Less. Position paper; research synthesis. https://upcube.ai/resources/research/better-models-increase-evaluation-value
  9. Ethen Research Lab (2026). Faros: Researching How Intelligence Should Choose Intelligence. Position paper. https://upcube.ai/resources/research/faros-research