All tests

Language

Arena instruction following

Blind votes on prompts with specific instructions, judged on how well each answer follows them.

Published by Arena under CC BY 4.0. Compare it with other tests

RSS feed
Top score1529Gemini 4 Argon High, Google1 other model within its margin
Models tested389from 25 makers
Median score1334Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
GoogleGemini 4 Argon High1529#1 of 389
AnthropicClaude Opus 5.5 (high)1517#2 of 389
OpenAIGPT-5.6 Sol (xhigh)1490#9 of 389
Moonshot AIKimi K3 (max)1488#11 of 389
MetaMuse Spark 1.3 (max)1486#13 of 389
DeepSeekDeepSeek V4.1 Flash (max)1479#18 of 389
Z.aiGLM-5.3 (max)1476#21 of 389
XiaomiMiMo-V2.6-Pro1474#23 of 389
AlibabaQwen3.8 Max1474#25 of 389
xAIGrok 4.51461#40 of 389
BaiduERNIE 5.11453#50 of 389
TencentHy31446#61 of 389

New top scores

Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.

  1. Gemini 4 Argon High took the top spot at 1529 from Claude Opus 5.5, which had led at 1516.
  2. Claude Opus 5.5 took the top spot at 1516 from Claude Opus 4.6, which had led at 1514.Within margin

Every result

389 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Gemini 4 Argon HighGoogle15291514 to 1543
#2Claude Opus 5.5 (high)Anthropic15171502 to 1533
#3Claude Opus 4.6 (high)Anthropic15141509 to 1519
#4Claude Fable 5 (high)Anthropic15091503 to 1515
#5Claude Opus 4.7 (high)Anthropic15021497 to 1508
#6Claude Fable 5.1 (max)Anthropic14981487 to 1508
#7Claude Opus 5 (high)Anthropic14961490 to 1502
#8Claude Opus 4.8 (high)Anthropic14911485 to 1496
#9GPT-5.6 Sol (xhigh)OpenAI14901484 to 1497
#10GPT-6.1 Sol (max)OpenAI14901471 to 1509
#11Kimi K3 (max)Moonshot AI14881481 to 1495
#12Gemini 3.8 Flash (high)Google14861479 to 1493
#13Muse Spark 1.3 (max)Meta14861476 to 1495
#14Gemini 3.7 Flash (high)Google14851477 to 1492
#15Claude Sonnet 5.5 (xhigh)Anthropic14841466 to 1502
#16Claude Opus 4.5 (20251101) (high 32k)Anthropic14831476 to 1490
#17Gemini 3.1 Pro PreviewGoogle14801475 to 1484
#18DeepSeek V4.1 Flash (max)DeepSeek14791469 to 1490
#19GPT-5.5 (high)OpenAI14781472 to 1483
#20Muse Spark 1.2 ( xhigh )Meta14771462 to 1492
#21GLM-5.3 (max)Z.ai14761467 to 1484
#22Claude Sonnet 4.6Anthropic14741469 to 1480
#23MiMo-V2.6-ProXiaomi14741458 to 1490
#24GPT-6 Astra (max)OpenAI14741463 to 1485
#25Qwen3.8 MaxAlibaba14741466 to 1481

How it is scored

Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.

Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.