Choosing a model
Guidance for selecting the right model for your task — balancing quality, speed, cost, and provider availability.
Start from the task, not the model. The fastest path to a good choice is to narrow by modality, quality floor, and latency budget — then pick the smallest model that clears those bars.
Decision framework
| Factor | Ask yourself | Example |
|---|---|---|
| Modality | Does the task need text, images, code, or audio? | Code generation needs a model strong in code. |
| Quality floor | How much does a wrong answer cost? | Production analytics need high accuracy; summarisation can tolerate more variance. |
| Latency budget | How fast does the response need to arrive? | Interactive chat needs sub-second; batch processing can wait. |
| Cost sensitivity | Is this high-volume or experimentation? | Prototyping can use premium models; production may need cheaper runners. |
| Provider reliability | Can you tolerate a single-provider outage? | Critical workloads should configure a fallback provider chain. |
Quick reference
| If you need... | Consider starting with... |
|---|---|
| General-purpose chat | Anthropic Claude Sonnet, OpenAI GPT-4o |
| Code generation | Anthropic Claude Sonnet, DeepSeek V3 |
| Fast, cheap completions | Meta Llama 3, Mistral, DeepSeek (smaller variants) |
| Image understanding | Anthropic Claude Sonnet, OpenAI GPT-4o |
| Local / offline inference | Llama 3, Mistral, Phi-3 (via Ollama or LM Studio) |
| Cost-sensitive production | Route to the cheapest model that meets quality, with premium fallback |
Using routing and fallback
The Gateway lets you set a model alias and a provider priority chain. When the primary provider is unavailable or returns an error, the Gateway retries the next provider in the chain automatically. This means you can:
- Use a cost-effective model for most requests.
- Fall back to a premium model when the primary fails.
- Avoid single-provider brittleness without client-side retry logic.
See the Gateway overview for implementation details.
Model Intelligence for deeper comparison
For data-driven comparisons — benchmark scores, latency percentiles, and pricing breakdowns — open the Model Intelligence portal. It provides per-model profiles, leaderboards, and provider-level aggregate statistics.
See Model Intelligence profiles for what is available and the freshness caveats.