All tests

Math

Arena math

Blind votes on the math questions people asked two anonymous models.

Published by Arena under CC BY 4.0. Compare it with other tests

RSS feed
Top score1530Gemini 4 Argon High, Google16 other models within its margin
Models tested372from 24 makers
Median score1351Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
GoogleGemini 4 Argon High1530#1 of 372
AnthropicClaude Fable 5 (high)1522#2 of 372
MetaMuse Spark 1.3 (max)1509#8 of 372
Z.aiGLM-5.3-Flash1503#11 of 372
DeepSeekDeepSeek V4.1 Flash (max)1503#12 of 372
Moonshot AIKimi K3 (max)1501#14 of 372
OpenAIGPT-5.51500#15 of 372
AlibabaQwen3.8 Max1498#16 of 372
XiaomiMiMo-V2.6-Pro1487#26 of 372
TencentHy31480#29 of 372
BaiduERNIE 5.11477#33 of 372
xAIGrok 4.51473#40 of 372

New top scores

Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.

  1. Gemini 4 Argon High took the top spot at 1528 from Claude Fable 5, which had led at 1523.Within margin

Every result

372 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Gemini 4 Argon HighGoogle15301495 to 1566
#2Claude Fable 5 (high)Anthropic15221509 to 1536
#3Claude Opus 5 (high)Anthropic15211509 to 1533
#4Claude Fable 5.1 (max)Anthropic15181493 to 1543
#5Gemini 3.8 Flash (high)Google15181501 to 1534
#6Claude Opus 4.6 (high)Anthropic15171507 to 1527
#7Claude Opus 5.5 (high)Anthropic15111475 to 1546
#8Muse Spark 1.3 (max)Meta15091485 to 1533
#9Claude Opus 4.7 (high)Anthropic15041493 to 1515
#10Gemini 3.7 Flash (high)Google15031485 to 1522
#11GLM-5.3-FlashZ.ai15031485 to 1521
#12DeepSeek V4.1 Flash (max)DeepSeek15031476 to 1530
#13Gemini 3.6 Flash (high)Google15011487 to 1516
#14Kimi K3 (max)Moonshot AI15011484 to 1517
#15GPT-5.5OpenAI15001489 to 1510
#16Qwen3.8 MaxAlibaba14981480 to 1515
#17Gemini 3.5 Flash (high)Google14971485 to 1510
#18GLM-5.3 (max)Z.ai14941473 to 1515
#19Claude Opus 4.8 (high)Anthropic14941483 to 1505
#20GPT-5.6 Sol (xhigh)OpenAI14941480 to 1508
#21Muse Spark 1.1Meta14921478 to 1507
#22GPT-5.4 (high)OpenAI14921481 to 1502
#23GPT-6 Astra (max)OpenAI14901463 to 1517
#24Gemini 3.1 Pro PreviewGoogle14881480 to 1496
#24GLM-5.2 (max)Z.ai14881475 to 1501

How it is scored

Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.

Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.