Beta

Choosing a model

Guidance for selecting the right model for your task — balancing quality, speed, cost, and provider availability.

Raw

Start from the task, not the model. The fastest path to a good choice is to narrow by modality, quality floor, and latency budget — then pick the smallest model that clears those bars.

Decision framework

FactorAsk yourselfExample
ModalityDoes the task need text, images, code, or audio?Code generation needs a model strong in code.
Quality floorHow much does a wrong answer cost?Production analytics need high accuracy; summarisation can tolerate more variance.
Latency budgetHow fast does the response need to arrive?Interactive chat needs sub-second; batch processing can wait.
Cost sensitivityIs this high-volume or experimentation?Prototyping can use premium models; production may need cheaper runners.
Provider reliabilityCan you tolerate a single-provider outage?Critical workloads should configure a fallback provider chain.

Quick reference

If you need...Consider starting with...
General-purpose chatAnthropic Claude Sonnet, OpenAI GPT-4o
Code generationAnthropic Claude Sonnet, DeepSeek V3
Fast, cheap completionsMeta Llama 3, Mistral, DeepSeek (smaller variants)
Image understandingAnthropic Claude Sonnet, OpenAI GPT-4o
Local / offline inferenceLlama 3, Mistral, Phi-3 (via Ollama or LM Studio)
Cost-sensitive productionRoute to the cheapest model that meets quality, with premium fallback

Using routing and fallback

The Gateway lets you set a model alias and a provider priority chain. When the primary provider is unavailable or returns an error, the Gateway retries the next provider in the chain automatically. This means you can:

  • Use a cost-effective model for most requests.
  • Fall back to a premium model when the primary fails.
  • Avoid single-provider brittleness without client-side retry logic.

See the Gateway overview for implementation details.

Model Intelligence for deeper comparison

For data-driven comparisons — benchmark scores, latency percentiles, and pricing breakdowns — open the Model Intelligence portal. It provides per-model profiles, leaderboards, and provider-level aggregate statistics.

See Model Intelligence profiles for what is available and the freshness caveats.

Last verified 2026-07-10 · Owner models-team