Coding
Arena WebDev
People describe a web page or app; two anonymous models each build it, and people vote for the better working result.
Published by Arena under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Claude Opus 5.5 (max) | 1815 | #1 of 123 | |
| GPT-6 Astra (max) | 1788 | #2 of 123 | |
| Gemini 4 Argon High | 1680 | #8 of 123 | |
| Qwen3.8 Max | 1671 | #9 of 123 | |
| Kimi K3 (max) | 1658 | #11 of 123 | |
| Muse Spark 1.3 (max) | 1657 | #12 of 123 | |
| Grok 4.7 (xhigh) | 1638 | #13 of 123 | |
| Hy4 preview | 1633 | #15 of 123 | |
| GLM-5.3 (max) | 1623 | #17 of 123 | |
| DeepSeek V4.1 Flash (max) | 1620 | #18 of 123 | |
| MiMo-V2.6-Pro | 1618 | #21 of 123 | |
| Step 5 Preview (high) | 1570 | #30 of 123 |
Every result
123 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Claude Opus 5.5 (max) | 18151799 to 1831 | |
| #2 | GPT-6 Astra (max) | 17881777 to 1798 | |
| #3 | Claude Sonnet 5.5 (xhigh) | 17861768 to 1804 | |
| #4 | GPT-6.1 Sol (max) | 17581741 to 1775 | |
| #5 | Claude Fable 5.1 (max) | 17491740 to 1759 | |
| #6 | Claude Opus 5 (max) | 16951689 to 1702 | |
| #7 | GPT-6 Sol (max) | 16891678 to 1700 | |
| #8 | Gemini 4 Argon High | 16801666 to 1693 | |
| #9 | Qwen3.8 Max | 16711659 to 1683 | |
| #10 | Qwen3.8 Max 0902 | 16701662 to 1678 | |
| #11 | Kimi K3 (max) | 16581651 to 1664 | |
| #12 | Muse Spark 1.3 (max) | 16571648 to 1665 | |
| #13 | Grok 4.7 (xhigh) | 16381626 to 1649 | |
| #13 | Qwen3.8 Flash Next | 16381629 to 1646 | |
| #15 | Hy4 preview | 16331624 to 1643 | |
| #16 | Claude Fable 5 (high) | 16251619 to 1632 | |
| #17 | GLM-5.3 (max) | 16231615 to 1631 | |
| #18 | DeepSeek V4.1 Flash (max) | 16201610 to 1630 | |
| #19 | gpt-5.6-sol-xhigh (codex-harness) | 16201614 to 1626 | |
| #20 | Grok 4.6 (high) | 16201612 to 1627 | |
| #21 | MiMo-V2.6-Pro | 16181606 to 1631 | |
| #22 | GLM-5.3-Flash | 16161608 to 1623 | |
| #23 | GLM-5.2 (max) | 16051599 to 1611 | |
| #24 | Gemini 3.7 Flash (high) | 15921584 to 1599 | |
| #25 | Qwen3.8 27B | 15911584 to 1598 |
How it is scored
Arena’s published WebDev scores. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.
Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.