Overall
Arena text
People ask two anonymous models their own question and vote for the better answer. The ranking comes from millions of these blind votes.
Published by Arena under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Gemini 4 Argon High | 1525 | #1 of 389 | |
| Claude Opus 4.6 (high) | 1505 | #2 of 389 | |
| Muse Spark 1.3 (max) | 1494 | #8 of 389 | |
| Kimi K3 (max) | 1488 | #13 of 389 | |
| GPT-5.6 Sol (xhigh) | 1484 | #17 of 389 | |
| Qwen3.8 Max | 1482 | #20 of 389 | |
| MiMo-V2.6-Pro | 1480 | #23 of 389 | |
| GLM-5.3 (max) | 1479 | #24 of 389 | |
| Grok 4.20 Beta1 | 1475 | #31 of 389 | |
| DeepSeek V4.1 Flash (max) | 1474 | #32 of 389 | |
| ERNIE 5.1 | 1468 | #41 of 389 | |
| Seed 2.0 Pro | 1457 | #55 of 389 |
New top scores
Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.
- Gemini 4 Argon High took the top spot at 1525 from Claude Opus 5.5, which had led at 1509.
- Claude Opus 5.5 took the top spot at 1509 from Claude Fable 5, which had led at 1506.Within margin
Every result
389 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Gemini 4 Argon High | 15251516 to 1534 | |
| #2 | Claude Opus 4.6 (high) | 15051501 to 1508 | |
| #3 | Claude Fable 5 (high) | 15041500 to 1509 | |
| #4 | Claude Opus 5.5 (high) | 15041495 to 1513 | |
| #5 | Claude Opus 4.7 (high) | 15011498 to 1505 | |
| #6 | Claude Fable 5.1 (max) | 15011495 to 1508 | |
| #7 | Gemini 3.8 Flash (high) | 14951490 to 1500 | |
| #8 | Muse Spark 1.3 (max) | 14941488 to 1501 | |
| #9 | Muse Spark 1.2 ( xhigh ) | 14941484 to 1503 | |
| #10 | Muse Spark 1.1 | 14911487 to 1496 | |
| #11 | Claude Opus 5 (high) | 14901486 to 1494 | |
| #12 | Muse Spark | 14891484 to 1495 | |
| #13 | Kimi K3 (max) | 14881484 to 1493 | |
| #14 | Gemini 3.7 Flash (high) | 14881483 to 1493 | |
| #15 | Gemini 3.1 Pro Preview | 14871484 to 1490 | |
| #16 | Gemini 3 Pro | 14861482 to 1489 | |
| #17 | GPT-5.6 Sol (xhigh) | 14841480 to 1488 | |
| #18 | GPT-6.1 Sol (max) | 14831473 to 1494 | |
| #19 | Gemini 3.6 Flash (high) | 14831479 to 1487 | |
| #20 | Qwen3.8 Max | 14821476 to 1487 | |
| #21 | Claude Opus 4.8 (high) | 14821478 to 1485 | |
| #22 | GPT-5.5 (high) | 14811478 to 1485 | |
| #23 | MiMo-V2.6-Pro | 14801471 to 1490 | |
| #24 | GLM-5.3 (max) | 14791473 to 1484 | |
| #25 | Gemini 3.5 Flash (high) | 14771474 to 1481 |
How it is scored
Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.
Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.