Leaderboards
Benchmarks
Each benchmark page explains the evaluation dimension, shows metric direction and units, and ranks models only when certified results exist.
Source: Canonical model registry at build time
Benchmarks
- 01Intelligence BenchmarkHigher is betterThe Intelligence Benchmark measures overall model capability using a composite of evaluations across reasoning, coding, agentic tasks, instruction following, and multimodal understanding. Scores are derived from normalized evaluations on the Ethen Intelligence Index. Higher scores indicate stronger general-purpose capability.0 models ranked
- 02Speed BenchmarkHigher is betterThe Speed Benchmark measures output token generation throughput in tokens per second. Speed is measured under standard conditions and reflects the model's raw generation efficiency. Higher throughput generally means faster responses for chat and generation-heavy workloads.0 models ranked
- 03Latency BenchmarkLower is betterThe Latency Benchmark measures time to first answer token (TTFT), representing how quickly a model begins generating after receiving a prompt. Lower latency is critical for interactive and real-time applications. Measured values include reasoning prep time where applicable.0 models ranked
- 04Pricing BenchmarkLower is betterThe Pricing Benchmark compares model input and output token prices. Input price is the cost per million tokens sent to the model; output price is the cost per million tokens generated. Lower prices mean more cost-effective operation at scale.0 models ranked
- 05Cost Efficiency BenchmarkLower is betterThe Cost Efficiency Benchmark evaluates models on combined input and output pricing. This is not a single numerical score but a comparative view across both pricing dimensions, helping identify models that offer the best value for different workload profiles.0 models ranked
- 06Context Window BenchmarkHigher is betterThe Context Window Benchmark compares the maximum context length each model supports. A larger context window allows the model to process longer documents, maintain extended conversation history, and handle complex multi-turn tasks without truncation.0 models ranked
Models ranked
No certified ranking evidence is published yet.
Source
Canonical model registry at build time