Math
Arena math
Blind votes on the math questions people asked two anonymous models.
Published by Arena under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Gemini 4 Argon High | 1530 | #1 of 372 | |
| Claude Fable 5 (high) | 1522 | #2 of 372 | |
| Muse Spark 1.3 (max) | 1509 | #8 of 372 | |
| GLM-5.3-Flash | 1503 | #11 of 372 | |
| DeepSeek V4.1 Flash (max) | 1503 | #12 of 372 | |
| Kimi K3 (max) | 1501 | #14 of 372 | |
| GPT-5.5 | 1500 | #15 of 372 | |
| Qwen3.8 Max | 1498 | #16 of 372 | |
| MiMo-V2.6-Pro | 1487 | #26 of 372 | |
| Hy3 | 1480 | #29 of 372 | |
| ERNIE 5.1 | 1477 | #33 of 372 | |
| Grok 4.5 | 1473 | #40 of 372 |
New top scores
Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.
- Gemini 4 Argon High took the top spot at 1528 from Claude Fable 5, which had led at 1523.Within margin
Every result
372 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Gemini 4 Argon High | 15301495 to 1566 | |
| #2 | Claude Fable 5 (high) | 15221509 to 1536 | |
| #3 | Claude Opus 5 (high) | 15211509 to 1533 | |
| #4 | Claude Fable 5.1 (max) | 15181493 to 1543 | |
| #5 | Gemini 3.8 Flash (high) | 15181501 to 1534 | |
| #6 | Claude Opus 4.6 (high) | 15171507 to 1527 | |
| #7 | Claude Opus 5.5 (high) | 15111475 to 1546 | |
| #8 | Muse Spark 1.3 (max) | 15091485 to 1533 | |
| #9 | Claude Opus 4.7 (high) | 15041493 to 1515 | |
| #10 | Gemini 3.7 Flash (high) | 15031485 to 1522 | |
| #11 | GLM-5.3-Flash | 15031485 to 1521 | |
| #12 | DeepSeek V4.1 Flash (max) | 15031476 to 1530 | |
| #13 | Gemini 3.6 Flash (high) | 15011487 to 1516 | |
| #14 | Kimi K3 (max) | 15011484 to 1517 | |
| #15 | GPT-5.5 | 15001489 to 1510 | |
| #16 | Qwen3.8 Max | 14981480 to 1515 | |
| #17 | Gemini 3.5 Flash (high) | 14971485 to 1510 | |
| #18 | GLM-5.3 (max) | 14941473 to 1515 | |
| #19 | Claude Opus 4.8 (high) | 14941483 to 1505 | |
| #20 | GPT-5.6 Sol (xhigh) | 14941480 to 1508 | |
| #21 | Muse Spark 1.1 | 14921478 to 1507 | |
| #22 | GPT-5.4 (high) | 14921481 to 1502 | |
| #23 | GPT-6 Astra (max) | 14901463 to 1517 | |
| #24 | Gemini 3.1 Pro Preview | 14881480 to 1496 | |
| #24 | GLM-5.2 (max) | 14881475 to 1501 |
How it is scored
Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.
Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.