Language
Arena instruction following
Blind votes on prompts with specific instructions, judged on how well each answer follows them.
Published by Arena under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Gemini 4 Argon High | 1529 | #1 of 389 | |
| Claude Opus 5.5 (high) | 1517 | #2 of 389 | |
| GPT-5.6 Sol (xhigh) | 1490 | #9 of 389 | |
| Kimi K3 (max) | 1488 | #11 of 389 | |
| Muse Spark 1.3 (max) | 1486 | #13 of 389 | |
| DeepSeek V4.1 Flash (max) | 1479 | #18 of 389 | |
| GLM-5.3 (max) | 1476 | #21 of 389 | |
| MiMo-V2.6-Pro | 1474 | #23 of 389 | |
| Qwen3.8 Max | 1474 | #25 of 389 | |
| Grok 4.5 | 1461 | #40 of 389 | |
| ERNIE 5.1 | 1453 | #50 of 389 | |
| Hy3 | 1446 | #61 of 389 |
New top scores
Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.
- Gemini 4 Argon High took the top spot at 1529 from Claude Opus 5.5, which had led at 1516.
- Claude Opus 5.5 took the top spot at 1516 from Claude Opus 4.6, which had led at 1514.Within margin
Every result
389 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Gemini 4 Argon High | 15291514 to 1543 | |
| #2 | Claude Opus 5.5 (high) | 15171502 to 1533 | |
| #3 | Claude Opus 4.6 (high) | 15141509 to 1519 | |
| #4 | Claude Fable 5 (high) | 15091503 to 1515 | |
| #5 | Claude Opus 4.7 (high) | 15021497 to 1508 | |
| #6 | Claude Fable 5.1 (max) | 14981487 to 1508 | |
| #7 | Claude Opus 5 (high) | 14961490 to 1502 | |
| #8 | Claude Opus 4.8 (high) | 14911485 to 1496 | |
| #9 | GPT-5.6 Sol (xhigh) | 14901484 to 1497 | |
| #10 | GPT-6.1 Sol (max) | 14901471 to 1509 | |
| #11 | Kimi K3 (max) | 14881481 to 1495 | |
| #12 | Gemini 3.8 Flash (high) | 14861479 to 1493 | |
| #13 | Muse Spark 1.3 (max) | 14861476 to 1495 | |
| #14 | Gemini 3.7 Flash (high) | 14851477 to 1492 | |
| #15 | Claude Sonnet 5.5 (xhigh) | 14841466 to 1502 | |
| #16 | Claude Opus 4.5 (20251101) (high 32k) | 14831476 to 1490 | |
| #17 | Gemini 3.1 Pro Preview | 14801475 to 1484 | |
| #18 | DeepSeek V4.1 Flash (max) | 14791469 to 1490 | |
| #19 | GPT-5.5 (high) | 14781472 to 1483 | |
| #20 | Muse Spark 1.2 ( xhigh ) | 14771462 to 1492 | |
| #21 | GLM-5.3 (max) | 14761467 to 1484 | |
| #22 | Claude Sonnet 4.6 | 14741469 to 1480 | |
| #23 | MiMo-V2.6-Pro | 14741458 to 1490 | |
| #24 | GPT-6 Astra (max) | 14741463 to 1485 | |
| #25 | Qwen3.8 Max | 14741466 to 1481 |
How it is scored
Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.
Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.