All tests

Overall

Arena text

People ask two anonymous models their own question and vote for the better answer. The ranking comes from millions of these blind votes.

Published by Arena under CC BY 4.0. Compare it with other tests

RSS feed
Top score1525Gemini 4 Argon High, GoogleNo other model within its margin
Models tested389from 25 makers
Median score1348Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
GoogleGemini 4 Argon High1525#1 of 389
AnthropicClaude Opus 4.6 (high)1505#2 of 389
MetaMuse Spark 1.3 (max)1494#8 of 389
Moonshot AIKimi K3 (max)1488#13 of 389
OpenAIGPT-5.6 Sol (xhigh)1484#17 of 389
AlibabaQwen3.8 Max1482#20 of 389
XiaomiMiMo-V2.6-Pro1480#23 of 389
Z.aiGLM-5.3 (max)1479#24 of 389
xAIGrok 4.20 Beta11475#31 of 389
DeepSeekDeepSeek V4.1 Flash (max)1474#32 of 389
BaiduERNIE 5.11468#41 of 389
ByteDanceSeed 2.0 Pro1457#55 of 389

New top scores

Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.

  1. Gemini 4 Argon High took the top spot at 1525 from Claude Opus 5.5, which had led at 1509.
  2. Claude Opus 5.5 took the top spot at 1509 from Claude Fable 5, which had led at 1506.Within margin

Every result

389 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Gemini 4 Argon HighGoogle15251516 to 1534
#2Claude Opus 4.6 (high)Anthropic15051501 to 1508
#3Claude Fable 5 (high)Anthropic15041500 to 1509
#4Claude Opus 5.5 (high)Anthropic15041495 to 1513
#5Claude Opus 4.7 (high)Anthropic15011498 to 1505
#6Claude Fable 5.1 (max)Anthropic15011495 to 1508
#7Gemini 3.8 Flash (high)Google14951490 to 1500
#8Muse Spark 1.3 (max)Meta14941488 to 1501
#9Muse Spark 1.2 ( xhigh )Meta14941484 to 1503
#10Muse Spark 1.1Meta14911487 to 1496
#11Claude Opus 5 (high)Anthropic14901486 to 1494
#12Muse SparkMeta14891484 to 1495
#13Kimi K3 (max)Moonshot AI14881484 to 1493
#14Gemini 3.7 Flash (high)Google14881483 to 1493
#15Gemini 3.1 Pro PreviewGoogle14871484 to 1490
#16Gemini 3 ProGoogle14861482 to 1489
#17GPT-5.6 Sol (xhigh)OpenAI14841480 to 1488
#18GPT-6.1 Sol (max)OpenAI14831473 to 1494
#19Gemini 3.6 Flash (high)Google14831479 to 1487
#20Qwen3.8 MaxAlibaba14821476 to 1487
#21Claude Opus 4.8 (high)Anthropic14821478 to 1485
#22GPT-5.5 (high)OpenAI14811478 to 1485
#23MiMo-V2.6-ProXiaomi14801471 to 1490
#24GLM-5.3 (max)Z.ai14791473 to 1484
#25Gemini 3.5 Flash (high)Google14771474 to 1481

How it is scored

Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.

Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.