LLM Leaderboard

Ranked models

Compare quality, speed, latency, and cost across available models.

#ModelOverall ScoreQuality (80%)Performance (20%)
Arena ELO40% weightArtificial Analysis Intelligence Index40% weightSpeed10% weightLatency10% weight
1
Claude Opus 5 icon
Claude Opus 5
Anthropic
0.903
1,493
50.8
127.0 t/s
2.71s
2
Muse Spark 1.3 icon
Muse Spark 1.3
Meta
0.867
1,493
48.1
78.0 t/s
3.80s
3
GPT-5.6 Sol icon
GPT-5.6 Sol
OpenAI
0.837
1,483
47.0
85.0 t/s
2.93s
4
Gemini 3.8 Flash icon
Gemini 3.8 Flash
Google
0.833
1,493
40.9
218.5 t/s
2.24s
5
Kimi K3 icon
Kimi K3
Moonshot
0.826
1,485
43.6
86.0 t/s
0.59s
6
GLM 5.3 Flash icon
GLM 5.3 Flash
Z.ai
0.785
1,475
41.8
94.0 t/s
0.76s
7
qwen3.8-max icon
qwen3.8-max
Alibaba
0.772
1,481
40.2
39.0 t/s
2.12s
8
GPT-5.6 Terra icon
GPT-5.6 Terra
OpenAI
0.765
1,466
42.1
101.0 t/s
0.69s
9
GPT-5.5 icon
GPT-5.5
OpenAI
0.747
1,482
38.4
97.0 t/s
5.71s
10
Grok 4.6 icon
Grok 4.6
SpaceXAI
0.722
1,456
44.3
99.5 t/s
6.95s
11
Gemini 3.1 Pro Preview icon
Gemini 3.1 Pro Preview
Google
0.703
1,487
29.7
103.0 t/s
2.96s
12
Claude Sonnet 5 icon
Claude Sonnet 5
Anthropic
0.702
1,461
38.2
71.0 t/s
2.93s
13
GPT-5.6 Luna icon
GPT-5.6 Luna
OpenAI
0.686
37.3
135.0 t/s
0.55s
14
GLM 5.2 icon
GLM 5.2
Z.ai
0.671
1,472
27.9
182.5 t/s
0.65s
15
Grok 4.20 icon
Grok 4.20
SpaceXAI
0.645
1,475
25.7
74.0 t/s
0.65s

About these metrics

Score

Composite score (0–1) blending the selected category’s quality metrics with speed and latency. Higher is better, relative to the models currently listed.

Quality (80%)

Arena ELO
Chatbot Arena ELO overall rating.40% weight
Artificial Analysis Intelligence Index
Measures general reasoning and knowledge capabilities.40% weight

Performance (20%)

Speed
Output throughput in tokens per second from the fastest available provider.10% weight
Latency
Time to first token in seconds — lower is better.10% weight

Use the + in the table header to view additional reference benchmarks and pricing metrics.

0 / 6 selected