Reasoning
Arena hard prompts
Blind votes on the hardest prompts: complex, specific questions that need real problem solving.
Published by Arena under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Gemini 4 Argon High | 1551 | #1 of 389 | |
| Claude Opus 5.5 (high) | 1534 | #2 of 389 | |
| Kimi K3 (max) | 1517 | #7 of 389 | |
| Muse Spark 1.3 (max) | 1517 | #8 of 389 | |
| GPT-5.6 Sol (xhigh) | 1511 | #14 of 389 | |
| MiMo-V2.6-Pro | 1506 | #19 of 389 | |
| GLM-5.3 (max) | 1503 | #21 of 389 | |
| Qwen3.8 Max | 1503 | #23 of 389 | |
| DeepSeek V4.1 Flash (max) | 1500 | #30 of 389 | |
| Grok 4.5 | 1489 | #42 of 389 | |
| ERNIE 5.1 | 1488 | #45 of 389 | |
| Step 5 Preview | 1482 | #53 of 389 |
New top scores
Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.
- Gemini 4 Argon High took the top spot at 1550 from Claude Opus 5.5, which had led at 1541.
- Claude Opus 5.5 took the top spot at 1541 from Claude Opus 4.6, which had led at 1533.Within margin
Every result
389 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Gemini 4 Argon High | 15511540 to 1562 | |
| #2 | Claude Opus 5.5 (high) | 15341523 to 1545 | |
| #3 | Claude Opus 4.6 (high) | 15331529 to 1537 | |
| #4 | Claude Fable 5 (high) | 15301525 to 1535 | |
| #5 | Claude Opus 4.7 (high) | 15261521 to 1530 | |
| #6 | Claude Fable 5.1 (max) | 15201513 to 1528 | |
| #7 | Kimi K3 (max) | 15171512 to 1523 | |
| #8 | Muse Spark 1.3 (max) | 15171510 to 1524 | |
| #9 | Gemini 3.8 Flash (high) | 15161510 to 1522 | |
| #10 | Claude Opus 5 (high) | 15151511 to 1520 | |
| #11 | Claude Opus 4.8 (high) | 15131509 to 1518 | |
| #12 | Muse Spark 1.1 | 15111506 to 1516 | |
| #13 | Muse Spark 1.2 ( xhigh ) | 15111500 to 1522 | |
| #14 | GPT-5.6 Sol (xhigh) | 15111505 to 1516 | |
| #15 | GPT-6.1 Sol (max) | 15091495 to 1523 | |
| #16 | Gemini 3.7 Flash (high) | 15081502 to 1514 | |
| #17 | Gemini 3.1 Pro Preview | 15071504 to 1511 | |
| #18 | Muse Spark | 15061500 to 1513 | |
| #19 | MiMo-V2.6-Pro | 15061494 to 1517 | |
| #20 | Claude Sonnet 5.5 (xhigh) | 15041491 to 1517 | |
| #21 | GLM-5.3 (max) | 15031497 to 1510 | |
| #22 | Claude Sonnet 4.6 | 15031499 to 1508 | |
| #23 | Qwen3.8 Max | 15031497 to 1509 | |
| #24 | Gemini 3 Pro | 15031498 to 1507 | |
| #25 | Gemini 3.6 Flash (high) | 15021497 to 1507 |
How it is scored
Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.
Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.