Language
Arena creative writing
Blind votes on the stories, poems and other writing people asked for.
Published by Arena under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Gemini 4 Argon High | 1519 | #1 of 387 | |
| Claude Opus 5.5 (high) | 1516 | #2 of 387 | |
| GPT-5.6 Sol (xhigh) | 1468 | #16 of 387 | |
| Qwen3.8 Max | 1468 | #17 of 387 | |
| Muse Spark | 1466 | #18 of 387 | |
| Grok 4.20 Beta1 | 1463 | #19 of 387 | |
| Kimi K3 (max) | 1460 | #20 of 387 | |
| GLM-5.2 (max) | 1454 | #26 of 387 | |
| MiMo-V2.6-Pro | 1446 | #38 of 387 | |
| DeepSeek V4 Pro | 1446 | #39 of 387 | |
| ERNIE 5.1 | 1428 | #63 of 387 | |
| Hy3 | 1422 | #68 of 387 |
New top scores
Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.
- Gemini 4 Argon High took the top spot at 1522 from Claude Opus 5.5, which had led at 1521.Within margin
- Claude Opus 5.5 took the top spot at 1521 from Claude Fable 5, which had led at 1504.Within margin
Every result
387 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Gemini 4 Argon High | 15191499 to 1538 | |
| #2 | Claude Opus 5.5 (high) | 15161497 to 1536 | |
| #3 | Claude Fable 5 (high) | 15031495 to 1510 | |
| #4 | Claude Opus 4.6 (high) | 15011494 to 1507 | |
| #5 | Gemini 3.7 Flash (high) | 14931484 to 1503 | |
| #6 | Claude Opus 4.7 (high) | 14891482 to 1496 | |
| #7 | Gemini 3.8 Flash (high) | 14851476 to 1494 | |
| #8 | Gemini 3 Pro | 14841476 to 1493 | |
| #9 | Claude Fable 5.1 (max) | 14821469 to 1494 | |
| #10 | Gemini 3.1 Pro Preview | 14801475 to 1486 | |
| #11 | Claude Opus 5 (high) | 14711465 to 1478 | |
| #12 | Gemini 3.6 Flash (high) | 14711463 to 1479 | |
| #13 | Claude Opus 4.5 (20251101) (high 32k) | 14701461 to 1478 | |
| #13 | Claude Opus 4.8 (high) | 14701463 to 1476 | |
| #15 | Gemini 3.5 Flash (medium) | 14691461 to 1476 | |
| #16 | GPT-5.6 Sol (xhigh) | 14681460 to 1476 | |
| #17 | Qwen3.8 Max | 14681458 to 1478 | |
| #18 | Muse Spark | 14661452 to 1479 | |
| #19 | Grok 4.20 Beta1 | 14631453 to 1473 | |
| #20 | Kimi K3 (max) | 14601452 to 1469 | |
| #21 | GPT-6.1 Sol (max) | 14601436 to 1483 | |
| #22 | Muse Spark 1.3 (max) | 14591446 to 1471 | |
| #23 | Gemini 3 Flash Preview | 14571448 to 1466 | |
| #24 | GPT-5.5 Instant | 14561447 to 1466 | |
| #25 | Muse Spark 1.2 ( xhigh ) | 14561435 to 1476 |
How it is scored
Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.
Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.