Position Paper · 2026-10-03 · Model Intelligence / Faros
Faros: Researching How Intelligence Should Choose Intelligence
Choosing a model is the smallest part of the decision. An agent system must also choose how much context to use, which tools to expose, how hard to verify, when to retry and when to stop. We study that decision as a whole.
Abstract
AI model routing, sending each request to the cheapest model that can handle it, has become a standard cost lever, and a growing literature studies how to learn it. We argue that routing understood this way is both too narrow and too fragile to be a durable research program. It is too narrow because, for agents that act, the decisions that determine cost and reliability include context allocation, tool exposure, verification level, retry versus reconciliation, escalation and budget pacing, not only model choice. It is too fragile because the margin between expensive and cheap models shrinks as prices fall and capabilities converge. We propose studying Faros, Ethen's decision layer, as an execution-configuration policy. Faros chooses among feasible configurations to maximize expected verified outcomes per unit cost, inside hard constraints it can never override. We describe the decision space and a staged path from strong rules to learned policies, with gates that rules may win. We identify what must be logged from the first decision for later learning to be valid. This is a position paper. It reports no Ethen routing measurements.
The routing question, and why it is incomplete
The research literature on routing among language models is substantial. FrugalGPT showed that cascading from cheaper to more expensive models could reduce cost while preserving quality on several tasks (Chen et al.). Hybrid routing trains a router to send queries to a small or large model depending on predicted difficulty (Ding et al.). AutoMix uses self-verification to decide when to escalate (Aggarwal et al.). RouteLLM learns routers from preference data (Ong et al.). RouterBench provides a standardized evaluation (Hu et al.).
Two recent results temper the enthusiasm. LLMRouterBench re-evaluated routing methods at scale: over 400,000 instances, 21 datasets and 33 models. While confirming strong complementarity among models, it found that many methods perform similarly under unified evaluation, and that several recent approaches, including commercial routers, fail to reliably outperform a simple baseline [EXTERNAL PRIMARY-SOURCE RESULT] (Li et al.). It also found a substantial gap to the oracle, driven mainly by failures to recall which model would have succeeded, and that careful curation of a few models beats larger ensembles. Separately, a study of routers trained to predict scalar performance scores identified routing collapse: as the cost budget increases, routers default to the most expensive model even when cheaper ones suffice. The authors attribute this to a mismatch between predicting scores and making discrete comparisons, and mitigate it with a router that learns rankings directly (Lai et al.).
These findings suggest a research stance rather than a product claim. The deeper question is not which model is best. It is: given a task, a mandate and a budget, which complete execution configuration is most likely to produce a verified outcome at acceptable cost and risk?
Reframing: an execution-configuration policy
We define Faros's output not as a model name but as a permitted execution configuration. Its inputs are the task's features, the constraints of the governing mandate, tenant policy (including data residency and contractual model restrictions), live provider health, the remaining budget and the current context state. Its outputs are:
- model tier and provider lane;
- reasoning effort;
- context allocation: what to include in full, what to summarize, what to exclude;
- the tool set exposed to the agent;
- the verification level required before the task can close;
- a retry, reconciliation and escalation plan;
- a cache strategy.
Figure 1 shows the structure.
Figure 1. Faros as a constrained decision layer. Hard constraints from the mandate, tenant policy and contracts define the feasible set first; only then does Faros optimize expected verified outcomes per unit cost within it. Faros advises; it never grants authority. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal (Faros).
The ordering is a principle. Hard constraints define the feasible set first and are never traded against cost. If a mandate forbids sending data outside a region, no saving justifies routing to a provider in another region. Within the feasible set, Faros optimizes expected verified successes per unit cost, subject to latency and risk limits. Faros advises; it does not authorize. It cannot grant permissions, waive approvals or disable evidence collection.
Figure 2 contrasts the decision space with that of a conventional router.
Figure 2. Model routing versus execution-configuration policy. What a conventional model router decides compared with the broader decision space proposed for Faros. The extra dimensions are where we hypothesize durable value lies; whether learning them beats rules is untested. Evidence label: QUALITATIVE MATRIX. Source: Ethen internal synthesis; qualitative.
Why the broader decision space matters
Four observations motivate the reframing.
Cost is dominated by the whole task, not one call. An agent task can involve dozens of model calls, tool calls, retries and verification steps. A cheap model that triggers two extra retries and a failed verification can cost more than an expensive model that succeeds once. The relevant objective is cost per verified outcome, defined in Cost Per Verified Outcome.
Context and tools are first-order levers. How much context to include, and how many tool schemas to expose, affect cost, latency and accuracy as much as model choice. Long contexts degrade some retrieval behaviors (Liu et al.), and context compaction can silently drop constraints (Wang et al.).
Verification is a decision. Whether to run a test suite, a second model's check or a human review is a cost–risk trade-off that depends on the task. Treating verification as fixed overhead hides one of the most important choices.
Switching has costs. Changing models mid-task can invalidate cached context, break consistency and confuse attribution of outcomes. Ethen's internal design work therefore proposes budgeting at the task level and keeping a model sticky within a chain of steps. Switches happen only at explicit boundaries, such as a hand-off to a verification step or a fresh-context subagent.
Cache economics versus decision intelligence
Some apparent routing gains are infrastructure effects. Prompt caching, batching and provider pricing changes can lower cost without any improvement in decisions. A router that happens to keep requests on a provider with a warm cache will look intelligent. Our research plan separates the two with factorial designs: rules versus learned policy, crossed with cache-aware versus cache-blind placement. This lets cache-economics gains be reported separately from decision-intelligence gains. The full protocol is How to Test Whether Learned AI Routing Beats Strong Rules.
Staged evolution, with rules as a valid end state
Figure 3 shows the proposed stages.
Figure 3. Staged evolution with promotion gates. Each stage must beat the one before on cost per verified outcome without a material quality loss. If a learned stage fails its gate, rules remain the production policy; rules are a legitimate end state. Evidence label: PROPOSED ARCHITECTURE. Source: Ethen architecture proposal; gates are PROPOSED TARGETS.
V1: strong rules. Capability, residency, provider health, cost, latency, context length, tool support, fallback and stickiness, encoded deterministically. V1 must be good enough to be a demanding baseline. It also logs, for every decision, the set of candidates considered and the probability with which each could have been selected.
V2: shadow decisions and bounded exploration. A learned policy makes decisions in shadow alongside V1. A small, mandate-bounded share of eligible traffic may be randomized among feasible options to collect unbiased comparisons. Contextual bandits with budget constraints are one natural mechanism here (Panda et al.). Off-policy estimates from logged data are checked against shadow results.
V3: learned ranking. A model predicts the probability of a verified outcome for each feasible configuration and ranks them, rather than predicting a scalar quality score. The routing-collapse result above is one reason to prefer ranking.
V4: configuration synthesis. Choosing harness policy, tools and verifier jointly rather than independently. This is a long-term research bet.
The promotion gate proposed in Ethen's internal reconciliation work is demanding by design [PROPOSED TARGET]. A learned stage must achieve at least 15% lower cost per verified outcome than strong rules, while the one-sided 95% lower confidence bound on the quality difference stays above −2 percentage points, with no degradation of safety or latency objectives. The comparison is against good rules, never against a frontier-only baseline, which simple tiering already beats. If a learned stage fails, the rules stay. That outcome is reported, not hidden. The survey Why Learned AI Model Routing Must Beat Good Rules develops the case for this bar.
What must be logged from day one
Learning a decision policy from logged decisions is a counterfactual problem: the log shows outcomes only for the configurations that were chosen. Unbiased off-policy evaluation requires knowing the probability with which each logged decision was made (Li et al., 2010; Dudík et al.; Swaminathan & Joachims). A deterministic rule assigns probability one to its choice and zero to every alternative, which leaves no data about alternatives at all. A system that does not log candidate sets and selection probabilities from its first decision accumulates data that no later volume can de-bias. This is the subject of Why AI Routers Should Log Propensities From Day One. The complementary source of counterfactuals, re-running completed tasks under alternative configurations in sandboxes, is described in Counterfactual Replay for AI Agents.
What survives cheaper models
If inference becomes ten times cheaper, the margin from sending easy requests to small models shrinks toward zero. We think the decision layer remains valuable, for a different reason: the knowledge of which configuration (context, tools, verification and recovery) produces verified outcomes on which work does not disappear with price changes. This is a hypothesis. A related analysis suggests that even cascading may have limited value in some settings: in a game-theoretic model of a provider routing between a standard and a reasoning model, the optimal policy was almost always static, with no cascading (Mahmood). Whether configuration knowledge transfers across model generations is measurable with the ledger described in The Capability Transfer Ledger. The same machinery also supports Model Change Assurance: predicting what changes when a tenant switches models.
Research questions
The program reduces to four questions, each with an answer that would change what we build:
- Separability. Are task classes in real agent workloads distinct enough that different configurations are optimal for different classes? If one configuration is best almost everywhere, a learned policy has little to learn.
- Learnability against rules. Given logged decisions with propensities, can a learned ranking beat strong rules on cost per verified outcome with quality protected? The gate above defines "beat".
- Decomposition. How much of any observed gain comes from cache placement and provider pricing rather than from better decisions? If almost all of it, the learned policy should not be credited.
- Durability. After a model upgrade, how much of a learned policy's advantage remains before re-fitting? If none, its value is a recurring cost of re-learning, not an accumulating asset.
None of these questions has been answered for Ethen workloads. All four can be answered with the logging and replay infrastructure described above.
Risks and counterarguments
- Rules may simply be good enough. That is a valid finding. The research is designed to discover it cheaply.
- Non-stationarity. Models, prices and provider quality change monthly. A learned policy is a depreciating asset that needs re-fitting and drift detection.
- Exploration has costs. Randomizing among feasible configurations, even within a mandate, sometimes chooses a worse option. Exploration budgets must be small, bounded by risk class and visible to customers where contracts require.
- Attribution is hard. Outcomes depend on many interacting choices. Factorial designs and careful ablations are expensive, and some interactions will remain unidentified.
Limitations
Faros, as described here, is an architecture and research agenda. No learned Faros stage has passed the proposed gate, and no Ethen routing result is reported. The proposed thresholds are hypotheses requiring power analysis. The claim that configuration knowledge outlasts routing margins is untested. External results cited here come from benchmarks whose task distributions may differ substantially from enterprise agent work.
Conclusion
The interesting question is not which model to call. It is how a system should spend its budget across models, context, tools, verification and recovery to produce work that can be verified, inside constraints it may not cross. Framed that way, routing becomes one dimension of a decision problem that remains meaningful as models get cheaper and better. We intend to study it with strong rules as the baseline, propensities logged from the first decision, and a willingness to conclude that rules win.
FAQ
Is Faros a model router? It includes model routing but is broader: it chooses a full execution configuration, including context, tools, verification and recovery, within hard constraints.
Why not train a learned router immediately? Because published evidence shows learned routers often fail to beat simple baselines. A learned policy should be promoted only when it beats strong rules on cost per verified outcome with quality protected.
Can Faros override a mandate to save money? No. Mandate, residency and policy constraints define the feasible set and are never optimized away.
Related research
- Why Learned AI Model Routing Must Beat Good Rules — why learned routing must beat good rules.
- Why AI Routers Should Log Propensities From Day One — propensity logging from the first decision.
- Counterfactual Replay for AI Agents — counterfactual data for configuration choice.
- How to Test Whether Learned AI Routing Beats Strong Rules — the protocol that decides V2/V3.
- Cost Per Verified Outcome: A Better Economic Unit for Agentic AI — CPVO as the objective.
- Mandates: Compiling Human Intent Into Bounded Agent Authority — mandates bound the decision space.
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work — routing machinery reused for upgrade assurance.
References
- Chen, L., Zaharia, M., Zou, J. (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176. https://arxiv.org/abs/2305.05176
- Ding, D. et al. (2024). Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing. arXiv:2404.14618. https://arxiv.org/abs/2404.14618
- Aggarwal, P. et al. (2023). AutoMix: Automatically Mixing Language Models. arXiv:2310.12963. https://arxiv.org/abs/2310.12963
- Ong, I. et al. (2024). RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665. https://arxiv.org/abs/2406.18665
- Hu, Q. J. et al. (2024). RouterBench: A Benchmark for Multi-LLM Routing System. arXiv:2403.12031. https://arxiv.org/abs/2403.12031
- Li, H. et al. (2026). LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. arXiv:2601.07206. https://arxiv.org/abs/2601.07206
- Lai, G. et al. (2026). When Routing Collapses: On the Degenerate Convergence of LLM Routers. arXiv:2602.03478. https://arxiv.org/abs/2602.03478
- Mahmood, R. (2026). Routing, Cascades, and User Choice for LLMs. arXiv:2602.09902. https://arxiv.org/abs/2602.09902
- Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172. https://arxiv.org/abs/2307.03172
- Wang, Z. et al. (2026). Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction. arXiv:2608.11242. https://arxiv.org/abs/2608.11242
- Li, L. et al. (2010). Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms. arXiv:1003.5956. https://arxiv.org/abs/1003.5956
- Dudík, M., Langford, J., Li, L. (2011). Doubly Robust Policy Evaluation and Learning. arXiv:1103.4601. https://arxiv.org/abs/1103.4601
- Swaminathan, A., Joachims, T. (2015). Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. arXiv:1502.02362. https://arxiv.org/abs/1502.02362
- Panda, P. et al. (2025). Adaptive LLM Routing under Budget Constraints. arXiv:2508.21141. https://arxiv.org/abs/2508.21141
More from Ethen Research Lab
Each publication states its evidence status. Designs, protocols, and proposals report no measured results.
- Why Learned AI Model Routing Must Beat Good Rules
A survey of learned LLM routing: what RouteLLM, RouterBench and LLMRouterBench show, why strong rules are the right baseline, and how to test non-inferiority.
- Model Change Assurance: Testing AI Upgrades Before They Reach Real Work
A research note on model change assurance: replaying an organization's own historical tasks to find regressions before an LLM upgrade reaches real work.
- How to Test Whether Learned AI Routing Beats Strong Rules
A research protocol for an LLM routing experiment: strong-rule baselines, propensity-logged data, factorial cache controls, non-inferiority tests and decision rules.
Explained on the Ethen Blog
- How Ethen Gateway Chooses an Eligible Model
Eligibility is elimination: the Gateway discards every model that cannot serve a request before it ranks anything at all.
- Choosing a Multi-Model AI Workspace
A workflow-first checklist for deciding when a general chat, a specialized app, a control plane, or local execution fits — using Ethen's stated product boundaries.
- How to Explore Ethen Research Lab: Programs, Evidence Labels and Reading Paths
The fastest way to read Ethen Research Lab well is to check two labels before reading anything else: the publication type (position paper, research note, proposal, technical report, methods paper, protocol, benchmark design, survey or system card) and the evidence status (measured result, synthesis, proposal, protocol, or external survey). Together they tell you what kind of claim the paper can make. Then filter the archive by research program to find papers on your topic, and use a reading path to follow a question from concept to benchmark to experiment. This guide explains each label, the programs, how related papers fit together, and where to start for your role.
Explore this topic
Related Ethen products
Ethen Research Lab is Upcube's public research publication program. It publishes papers, protocols, benchmark designs, and system cards, each labelled with its evidence status. It is separate from Ethen Research, the AI research workspace product.