All tests

Reasoning

Arena hard prompts

Blind votes on the hardest prompts: complex, specific questions that need real problem solving.

Published by Arena under CC BY 4.0. Compare it with other tests

RSS feed
Top score1551Gemini 4 Argon High, GoogleNo other model within its margin
Models tested389from 25 makers
Median score1363Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
GoogleGemini 4 Argon High1551#1 of 389
AnthropicClaude Opus 5.5 (high)1534#2 of 389
Moonshot AIKimi K3 (max)1517#7 of 389
MetaMuse Spark 1.3 (max)1517#8 of 389
OpenAIGPT-5.6 Sol (xhigh)1511#14 of 389
XiaomiMiMo-V2.6-Pro1506#19 of 389
Z.aiGLM-5.3 (max)1503#21 of 389
AlibabaQwen3.8 Max1503#23 of 389
DeepSeekDeepSeek V4.1 Flash (max)1500#30 of 389
xAIGrok 4.51489#42 of 389
BaiduERNIE 5.11488#45 of 389
StepFunStep 5 Preview1482#53 of 389

New top scores

Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.

  1. Gemini 4 Argon High took the top spot at 1550 from Claude Opus 5.5, which had led at 1541.
  2. Claude Opus 5.5 took the top spot at 1541 from Claude Opus 4.6, which had led at 1533.Within margin

Every result

389 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Gemini 4 Argon HighGoogle15511540 to 1562
#2Claude Opus 5.5 (high)Anthropic15341523 to 1545
#3Claude Opus 4.6 (high)Anthropic15331529 to 1537
#4Claude Fable 5 (high)Anthropic15301525 to 1535
#5Claude Opus 4.7 (high)Anthropic15261521 to 1530
#6Claude Fable 5.1 (max)Anthropic15201513 to 1528
#7Kimi K3 (max)Moonshot AI15171512 to 1523
#8Muse Spark 1.3 (max)Meta15171510 to 1524
#9Gemini 3.8 Flash (high)Google15161510 to 1522
#10Claude Opus 5 (high)Anthropic15151511 to 1520
#11Claude Opus 4.8 (high)Anthropic15131509 to 1518
#12Muse Spark 1.1Meta15111506 to 1516
#13Muse Spark 1.2 ( xhigh )Meta15111500 to 1522
#14GPT-5.6 Sol (xhigh)OpenAI15111505 to 1516
#15GPT-6.1 Sol (max)OpenAI15091495 to 1523
#16Gemini 3.7 Flash (high)Google15081502 to 1514
#17Gemini 3.1 Pro PreviewGoogle15071504 to 1511
#18Muse SparkMeta15061500 to 1513
#19MiMo-V2.6-ProXiaomi15061494 to 1517
#20Claude Sonnet 5.5 (xhigh)Anthropic15041491 to 1517
#21GLM-5.3 (max)Z.ai15031497 to 1510
#22Claude Sonnet 4.6Anthropic15031499 to 1508
#23Qwen3.8 MaxAlibaba15031497 to 1509
#24Gemini 3 ProGoogle15031498 to 1507
#25Gemini 3.6 Flash (high)Google15021497 to 1507

How it is scored

Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.

Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.