All tests

Coding

Arena coding

Blind votes on conversations about code: people compare two anonymous answers to their own programming question.

Published by Arena under CC BY 4.0. Compare it with other tests

RSS feed
Top score1560Gemini 4 Argon High, Google4 other models within its margin
Models tested384from 25 makers
Median score1392Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
GoogleGemini 4 Argon High1560#1 of 384
AnthropicClaude Fable 5 (high)1552#2 of 384
OpenAIGPT-6 Astra (max)1543#5 of 384
Moonshot AIKimi K3 (max)1541#7 of 384
XiaomiMiMo-V2.6-Pro1540#8 of 384
MetaMuse Spark 1.3 (max)1539#9 of 384
DeepSeekDeepSeek V4.1 Flash (max)1528#22 of 384
AlibabaQwen3.7 Max Preview1524#23 of 384
Z.aiGLM-5.3-Flash1523#25 of 384
xAIGrok 4.51514#41 of 384
ByteDanceSeed 2.0 Pro1514#42 of 384
BaiduERNIE 5.11512#47 of 384

New top scores

Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.

  1. Gemini 4 Argon High took the top spot at 1559 from Claude Opus 4.6, which had led at 1551.Within margin
  2. Claude Opus 4.6 took the top spot at 1551 from Claude Fable 5, which had led at 1552.Within margin

Every result

384 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Gemini 4 Argon HighGoogle15601543 to 1577
#2Claude Fable 5 (high)Anthropic15521545 to 1559
#3Claude Opus 4.6 (high)Anthropic15511546 to 1557
#4Claude Opus 4.7 (high)Anthropic15511545 to 1557
#5GPT-6 Astra (max)OpenAI15431530 to 1556
#6GPT-6.1 Sol (max)OpenAI15421518 to 1566
#7Kimi K3 (max)Moonshot AI15411534 to 1549
#8MiMo-V2.6-ProXiaomi15401522 to 1558
#9Muse Spark 1.3 (max)Meta15391528 to 1549
#10Claude Opus 5.5 (high)Anthropic15391521 to 1556
#11Claude Sonnet 5.5 (xhigh)Anthropic15361515 to 1558
#12GPT-5.6 Sol (xhigh)OpenAI15341527 to 1541
#13Claude Opus 5 (high)Anthropic15331527 to 1540
#14Claude Opus 4.8 (high)Anthropic15331527 to 1539
#15Muse Spark 1.1Meta15331526 to 1539
#16Muse Spark 1.2 ( xhigh )Meta15311514 to 1549
#17Claude Fable 5.1 (max)Anthropic15311519 to 1544
#18Claude Opus 4.5 (20251101) (high 32k)Anthropic15301523 to 1538
#19Gemini 3.8 Flash (high)Google15301522 to 1538
#19Muse SparkMeta15301520 to 1540
#21Claude Sonnet 4.6Anthropic15291523 to 1534
#22DeepSeek V4.1 Flash (max)DeepSeek15281516 to 1541
#23Qwen3.7 Max PreviewAlibaba15241506 to 1542
#24Qwen3.8 MaxAlibaba15241516 to 1532
#25GLM-5.3-FlashZ.ai15231515 to 1532

How it is scored

Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.

Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.