LLM Leaderboard
Ranked models
Compare quality, speed, latency, and cost across available models.
Synced Aug 6, 4:02 PM
| # | Model | Overall Score↓ | Quality (80%) | Performance (20%) | |||
|---|---|---|---|---|---|---|---|
| Arena ELO40% weight | Artificial Analysis Intelligence Index40% weight | Speed10% weight | Latency10% weight | ||||
| 1 | Claude Opus 5 Anthropic | 0.898 | 1,492 | 60.7 | 94.0 t/s | 2.40s | |
| 2 | Kimi K3 Moonshot | 0.857 | 1,485 | 57.1 | 98.0 t/s | 2.06s | |
| 3 | Qwen3.8 Max Alibaba | 0.856 | 1,496 | 56.2 | 37.0 t/s | 3.17s | |
| 4 | GPT-5.6 Sol OpenAI | 0.840 | 1,483 | 58.9 | 72.0 t/s | 4.97s | |
| 5 | GPT-5.5 OpenAI | 0.820 | 1,482 | 54.8 | 97.5 t/s | 4.09s | |
| 6 | Gemini 3.6 Flash Google | 0.801 | 1,483 | 50.1 | 92.0 t/s | 2.00s | |
| 7 | Grok 4.5 SpaceXAI | 0.787 | 1,469 | 53.8 | 54.0 t/s | 1.02s | |
| 8 | GPT-5.6 Luna OpenAI | 0.784 | — | 51.2 | 105.0 t/s | 0.95s | |
| 9 | GPT-5.6 Terra OpenAI | 0.779 | 1,468 | 55.0 | 94.5 t/s | 4.90s | |
| 10 | Claude Sonnet 5 Anthropic | 0.764 | 1,462 | 53.4 | 95.0 t/s | 3.08s | |
| 11 | Gemini 3.1 Pro Preview Google | 0.757 | 1,485 | 46.5 | 72.0 t/s | 4.91s | |
| 12 | Muse Spark 1.1 Meta | 0.729 | 1,490 | 50.6 | 78.0 t/s | 16.28s | |
| 13 | MiMo-V2.5-Pro Xiaomi | 0.707 | 1,466 | 42.2 | 84.0 t/s | 0.49s | |
| 14 | DeepSeek V4 Pro DeepSeek | 0.705 | 1,458 | 44.3 | 97.0 t/s | 0.45s | |
| 15 | GLM 5.2 Z.ai | 0.687 | 1,469 | 39.5 | 98.0 t/s | 2.45s | |
About these metrics
Score
Composite score (0–1) blending the selected category’s quality metrics with speed and latency. Higher is better, relative to the models currently listed.
Quality (80%)
- Arena ELO
- Chatbot Arena ELO overall rating.40% weight
- Artificial Analysis Intelligence Index
- Measures general reasoning and knowledge capabilities.40% weight
Performance (20%)
- Speed
- Output throughput in tokens per second from the fastest available provider.10% weight
- Latency
- Time to first token in seconds — lower is better.10% weight
Use the + in the table header to view additional reference benchmarks and pricing metrics.
0 / 6 selected