Coding
LiveBench coding
Code generation and completion on the published LiveBench tasks.
Published by LiveBench under Apache-2.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Claude Sonnet 5.5 (max effort) | 91.4% | #1 of 63 | |
| GPT-5.6 Sol (max) | 83.9% | #5 of 63 | |
| Kimi K3 | 81.5% | #13 of 63 | |
| DeepSeek V4.1 Flash (max) | 80.0% | #19 of 63 | |
| GLM-5.2 | 79.7% | #20 of 63 | |
| Gemini 3.7 Flash (high) | 78.9% | #26 of 63 | |
| Qwen3.6 Plus | 78.2% | #29 of 63 | |
| Grok 4.7 (xhigh) | 77.2% | #35 of 63 | |
| Nemotron 3 Ultra 550B A55B | 70.7% | #56 of 63 | |
| MiniMax-M3 | 68.2% | #61 of 63 |
New top scores
Each time a different model took the top spot, since the site began keeping this record.
- Claude Sonnet 5.5 took the top spot at 91.4% from Claude Opus 5.5, which had led at 89.3%.
- Claude Opus 5.5 took the top spot at 89.3% from Claude Fable 5.1, which had led at 86.4%.
Every result
63 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Claude Sonnet 5.5 (max effort) | 91.4% | |
| #2 | Claude Opus 5.5 (max effort) | 89.3% | |
| #3 | Claude Fable 5.1 (max effort) | 86.4% | |
| #4 | Claude Fable 5 (max effort) | 86.0% | |
| #5 | GPT-5.6 Sol (max) | 83.9% | |
| #6 | GPT-5.2 Codex | 83.6% | |
| #7 | GPT-5.6 Luna (max) | 82.9% | |
| #8 | Smaug Agentic | 82.5% | |
| #9 | GPT-5.5 (xhigh) | 82.2% | |
| #10 | Claude Opus 4.7 (xhigh effort) | 82.1% | |
| #11 | Claude Opus 4.8 (max effort) | 81.8% | |
| #12 | GPT-6 Sol (max) | 81.8% | |
| #13 | Claude Opus 5 (max effort) | 81.5% | |
| #13 | Kimi K3 | 81.5% | |
| #15 | Muse Spark 1.3 (xhigh) | 81.1% | |
| #16 | GPT-6.1 Sol (xhigh) | 80.7% | |
| #17 | Claude Sonnet 5 (xhigh effort) | 80.7% | |
| #18 | GPT-6 Astra (max) | 80.4% | |
| #19 | DeepSeek V4.1 Flash (max) | 80.0% | |
| #20 | Claude Opus 4.5 (20251101) (thinking 64k high effort) | 79.7% | |
| #20 | GLM-5.2 | 79.7% | |
| #22 | Claude Sonnet 4.6 (thinking auto medium effort) | 79.3% | |
| #23 | GLM-5.3 | 79.0% | |
| #23 | GLM-5.3-Flash | 79.0% | |
| #23 | GPT-6 Luna (max) | 79.0% |
How it is scored
Arithmetic mean of the 2 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.
Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.