All tests

Coding

Arena WebDev

People describe a web page or app; two anonymous models each build it, and people vote for the better working result.

Published by Arena under CC BY 4.0. Compare it with other tests

RSS feed
Top score1815Claude Opus 5.5 (max), AnthropicNo other model within its margin
Models tested123from 19 makers
Median score1439Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Opus 5.5 (max)1815#1 of 123
OpenAIGPT-6 Astra (max)1788#2 of 123
GoogleGemini 4 Argon High1680#8 of 123
AlibabaQwen3.8 Max1671#9 of 123
Moonshot AIKimi K3 (max)1658#11 of 123
MetaMuse Spark 1.3 (max)1657#12 of 123
xAIGrok 4.7 (xhigh)1638#13 of 123
TencentHy4 preview1633#15 of 123
Z.aiGLM-5.3 (max)1623#17 of 123
DeepSeekDeepSeek V4.1 Flash (max)1620#18 of 123
XiaomiMiMo-V2.6-Pro1618#21 of 123
StepFunStep 5 Preview (high)1570#30 of 123

Every result

123 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Opus 5.5 (max)Anthropic18151799 to 1831
#2GPT-6 Astra (max)OpenAI17881777 to 1798
#3Claude Sonnet 5.5 (xhigh)Anthropic17861768 to 1804
#4GPT-6.1 Sol (max)OpenAI17581741 to 1775
#5Claude Fable 5.1 (max)Anthropic17491740 to 1759
#6Claude Opus 5 (max)Anthropic16951689 to 1702
#7GPT-6 Sol (max)OpenAI16891678 to 1700
#8Gemini 4 Argon HighGoogle16801666 to 1693
#9Qwen3.8 MaxAlibaba16711659 to 1683
#10Qwen3.8 Max 0902Alibaba16701662 to 1678
#11Kimi K3 (max)Moonshot AI16581651 to 1664
#12Muse Spark 1.3 (max)Meta16571648 to 1665
#13Grok 4.7 (xhigh)xAI16381626 to 1649
#13Qwen3.8 Flash NextAlibaba16381629 to 1646
#15Hy4 previewTencent16331624 to 1643
#16Claude Fable 5 (high)Anthropic16251619 to 1632
#17GLM-5.3 (max)Z.ai16231615 to 1631
#18DeepSeek V4.1 Flash (max)DeepSeek16201610 to 1630
#19gpt-5.6-sol-xhigh (codex-harness)OpenAI16201614 to 1626
#20Grok 4.6 (high)xAI16201612 to 1627
#21MiMo-V2.6-ProXiaomi16181606 to 1631
#22GLM-5.3-FlashZ.ai16161608 to 1623
#23GLM-5.2 (max)Z.ai16051599 to 1611
#24Gemini 3.7 Flash (high)Google15921584 to 1599
#25Qwen3.8 27BAlibaba15911584 to 1598

How it is scored

Arena’s published WebDev scores. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.

Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.