Coding
Arena coding
Blind votes on conversations about code: people compare two anonymous answers to their own programming question.
Published by Arena under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Gemini 4 Argon High | 1560 | #1 of 384 | |
| Claude Fable 5 (high) | 1552 | #2 of 384 | |
| GPT-6 Astra (max) | 1543 | #5 of 384 | |
| Kimi K3 (max) | 1541 | #7 of 384 | |
| MiMo-V2.6-Pro | 1540 | #8 of 384 | |
| Muse Spark 1.3 (max) | 1539 | #9 of 384 | |
| DeepSeek V4.1 Flash (max) | 1528 | #22 of 384 | |
| Qwen3.7 Max Preview | 1524 | #23 of 384 | |
| GLM-5.3-Flash | 1523 | #25 of 384 | |
| Grok 4.5 | 1514 | #41 of 384 | |
| Seed 2.0 Pro | 1514 | #42 of 384 | |
| ERNIE 5.1 | 1512 | #47 of 384 |
New top scores
Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.
- Gemini 4 Argon High took the top spot at 1559 from Claude Opus 4.6, which had led at 1551.Within margin
- Claude Opus 4.6 took the top spot at 1551 from Claude Fable 5, which had led at 1552.Within margin
Every result
384 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Gemini 4 Argon High | 15601543 to 1577 | |
| #2 | Claude Fable 5 (high) | 15521545 to 1559 | |
| #3 | Claude Opus 4.6 (high) | 15511546 to 1557 | |
| #4 | Claude Opus 4.7 (high) | 15511545 to 1557 | |
| #5 | GPT-6 Astra (max) | 15431530 to 1556 | |
| #6 | GPT-6.1 Sol (max) | 15421518 to 1566 | |
| #7 | Kimi K3 (max) | 15411534 to 1549 | |
| #8 | MiMo-V2.6-Pro | 15401522 to 1558 | |
| #9 | Muse Spark 1.3 (max) | 15391528 to 1549 | |
| #10 | Claude Opus 5.5 (high) | 15391521 to 1556 | |
| #11 | Claude Sonnet 5.5 (xhigh) | 15361515 to 1558 | |
| #12 | GPT-5.6 Sol (xhigh) | 15341527 to 1541 | |
| #13 | Claude Opus 5 (high) | 15331527 to 1540 | |
| #14 | Claude Opus 4.8 (high) | 15331527 to 1539 | |
| #15 | Muse Spark 1.1 | 15331526 to 1539 | |
| #16 | Muse Spark 1.2 ( xhigh ) | 15311514 to 1549 | |
| #17 | Claude Fable 5.1 (max) | 15311519 to 1544 | |
| #18 | Claude Opus 4.5 (20251101) (high 32k) | 15301523 to 1538 | |
| #19 | Gemini 3.8 Flash (high) | 15301522 to 1538 | |
| #19 | Muse Spark | 15301520 to 1540 | |
| #21 | Claude Sonnet 4.6 | 15291523 to 1534 | |
| #22 | DeepSeek V4.1 Flash (max) | 15281516 to 1541 | |
| #23 | Qwen3.7 Max Preview | 15241506 to 1542 | |
| #24 | Qwen3.8 Max | 15241516 to 1532 | |
| #25 | GLM-5.3-Flash | 15231515 to 1532 |
How it is scored
Arena’s published scores with style control, which separates what an answer says from its length and formatting, as Arena’s leaderboard shows by default. Higher is better on an open-ended scale, like chess ratings; the range is Arena’s 95% confidence interval.
Results from Arena (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.