Flagships
LiveChoose models with evidence in view.
Compare capabilities, benchmarks, pricing, and evaluations where data is available.
Purpose
Compare for the work you need to do.
Consider capability, cost, latency, context, and modality where information is available. Evidence coverage varies by model.
Highlights
The leading models, three ways.
View as table
| Model | Provider | Intelligence | Speed (tokens/s) | Input price (USD / 1M) |
|---|---|---|---|---|
| Claude Fable 5 (with fallback) | Anthropic | 60 | 70.7 | $10.00 |
| Claude Opus 4.8 (max) | Anthropic | 56 | 64.5 | $5.00 |
| GPT-5.5 (xhigh) | OpenAI | 55 | 88 | $5.00 |
| Claude Opus 4.7 (max) | Anthropic | 54 | 56.4 | $5.00 |
| Claude Sonnet 5 (max) | Anthropic | 53 | 83.7 | $2.00 |
| GPT-5.5 (high) | OpenAI | 53 | 83.1 | $5.00 |
| GLM-5.2 (max) | Z AI | 51 | 217.6 | $1.40 |
| GPT-5.4 (xhigh) | OpenAI | 51 | 175.3 | $2.50 |
| Gemini 3.5 Flash | 50 | 192.2 | $1.50 | |
| GPT-5.5 (medium) | OpenAI | 50 | 72.7 | $5.00 |
The ten highest intelligence scores among 548 models in the Model Library. Source: Artificial Analysis scrape, normalized 2026-07-08. Third-party measurements, not Ethen evaluations.
Evidence
Keep the evidence types distinct.
Model facts, benchmark results, and Ethen evaluations answer different questions. Runtime availability describes access; production outcomes describe real use and should appear only where recorded and exposed.
| Dimension | Direction and unit | Models ranked |
|---|---|---|
| Intelligence | Higher · score | |
| Speed | Higher · tok/s | |
| Latency | Lower · s TTFT | |
| Input Price | Lower · $/1M tokens | |
| Output Price | Lower · $/1M tokens | |
| Context Window | Higher · tokens |
No certified ranking evidence is published yet.
Rankings are currently unavailable. Canonical results have not been certified.
Catalog And Access
Understand the model and its context.
Model Library is the catalog. Flagships supports comparison. A provider is where or how a model is served; a comparison profile alone does not establish hosting or runtime access.
Surfaces
Explore the public comparison portal.
Browse leaderboards on Ethen's public site. The authenticated product has a separate home on Platform.
FAQ
Do Flagships rank every model objectively?
Coverage varies, and a test result reflects its conditions. Use the available evidence and evaluate models for your own task.
Can I run every model I compare?
Execution depends on supported provider access or a compatible local runtime.
Find evidence for your next model choice.
Browse public comparisons and explore the catalog.