All tests

Language

Arena creative writing

Blind votes on the stories, poems and other writing people asked for.

Published by Arena under CC BY 4.0. Compare it with other tests

RSS feed
Top score1519Gemini 4 Argon High, Google3 other models within its margin
Models tested387from 25 makers
Median score1312Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
GoogleGemini 4 Argon High1519#1 of 387
AnthropicClaude Opus 5.5 (high)1516#2 of 387
OpenAIGPT-5.6 Sol (xhigh)1468#16 of 387
AlibabaQwen3.8 Max1468#17 of 387
MetaMuse Spark1466#18 of 387
xAIGrok 4.20 Beta11463#19 of 387
Moonshot AIKimi K3 (max)1460#20 of 387
Z.aiGLM-5.2 (max)1454#26 of 387
XiaomiMiMo-V2.6-Pro1446#38 of 387
DeepSeekDeepSeek V4 Pro1446#39 of 387
BaiduERNIE 5.11428#63 of 387
TencentHy31422#68 of 387

New top scores

Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.

  1. Gemini 4 Argon High took the top spot at 1522 from Claude Opus 5.5, which had led at 1521.Within margin
  2. Claude Opus 5.5 took the top spot at 1521 from Claude Fable 5, which had led at 1504.Within margin

Every result

387 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Gemini 4 Argon HighGoogle15191499 to 1538
#2Claude Opus 5.5 (high)Anthropic15161497 to 1536
#3Claude Fable 5 (high)Anthropic15031495 to 1510
#4Claude Opus 4.6 (high)Anthropic15011494 to 1507
#5Gemini 3.7 Flash (high)Google14931484 to 1503
#6Claude Opus 4.7 (high)Anthropic14891482 to 1496
#7Gemini 3.8 Flash (high)Google14851476 to 1494
#8Gemini 3 ProGoogle14841476 to 1493
#9Claude Fable 5.1 (max)Anthropic14821469 to 1494
#10Gemini 3.1 Pro PreviewGoogle14801475 to 1486
#11Claude Opus 5 (high)Anthropic14711465 to 1478
#12Gemini 3.6 Flash (high)Google14711463 to 1479
#13Claude Opus 4.5 (20251101) (high 32k)Anthropic14701461 to 1478
#13Claude Opus 4.8 (high)Anthropic14701463 to 1476
#15Gemini 3.5 Flash (medium)Google14691461 to 1476
#16GPT-5.6 Sol (xhigh)OpenAI14681460 to 1476
#17Qwen3.8 MaxAlibaba14681458 to 1478
#18Muse SparkMeta14661452 to 1479
#19Grok 4.20 Beta1xAI14631453 to 1473
#20Kimi K3 (max)Moonshot AI14601452 to 1469
#21GPT-6.1 Sol (max)OpenAI14601436 to 1483
#22Muse Spark 1.3 (max)Meta14591446 to 1471
#23Gemini 3 Flash PreviewGoogle14571448 to 1466
#24GPT-5.5 InstantOpenAI14561447 to 1466
#25Muse Spark 1.2 ( xhigh )Meta14561435 to 1476

How it is scored

Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.

Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.