All tests
Agents
LiveBench agentic coding
JavaScript, TypeScript and Python agentic coding tasks.
Published by LiveBench under Apache-2.0. Compare it with other tests
Top score77.3%DeepSeek V4.1 Flash (max), DeepSeek
Models tested63from 11 makers
Median score53.8%Half the models score above it
Scale0 to 100%The share of tasks done right
Top 15
Each model at its best configuration, coloured by maker.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| DeepSeek V4.1 Flash (max) | 77.3% | #1 of 63 | |
| Claude Opus 5.5 (max effort) | 71.7% | #2 of 63 | |
| Qwen3.8 Max | 64.7% | #6 of 63 | |
| Kimi K3 | 62.2% | #9 of 63 | |
| GLM-5.3 | 60.9% | #14 of 63 | |
| Gemini 3.7 Flash (high) | 58.3% | #18 of 63 | |
| GPT-6 Astra (max) | 57.3% | #20 of 63 | |
| Grok 4.6 | 57.0% | #21 of 63 | |
| MiniMax-M3 | 40.7% | #58 of 63 | |
| Nemotron 3 Ultra 550B A55B | 38.7% | #61 of 63 |
Every result
63 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | DeepSeek V4.1 Flash (max) | 77.3% | |
| #2 | Claude Opus 5.5 (max effort) | 71.7% | |
| #3 | Claude Fable 5.1 (max effort) | 66.1% | |
| #4 | Claude Opus 5 (max effort) | 65.2% | |
| #5 | DeepSeek V4 Flash Vision Exp | 65.1% | |
| #6 | Qwen3.8 Max | 64.7% | |
| #6 | Smaug Agentic | 64.7% | |
| #8 | Muse Spark 1.3 (xhigh) | 64.1% | |
| #9 | Claude Fable 5 (max effort) | 62.2% | |
| #9 | Kimi K3 | 62.2% | |
| #11 | Qwen3.8 Flash Next | 61.6% | |
| #12 | Qwen3.8 27B | 61.4% | |
| #13 | Smaug Flash | 61.1% | |
| #14 | GLM-5.3 | 60.9% | |
| #15 | Smaug Mini | 60.8% | |
| #16 | Claude Sonnet 5 (xhigh effort) | 59.4% | |
| #17 | Muse Spark 1.1 (xhigh) | 58.5% | |
| #18 | Gemini 3.7 Flash (high) | 58.3% | |
| #19 | Muse Spark 1.2 (xhigh) | 57.6% | |
| #20 | GPT-6 Astra (max) | 57.3% | |
| #21 | Grok 4.6 | 57.0% | |
| #22 | GLM-5.3-Flash | 56.8% | |
| #22 | GPT-6.1 Sol (xhigh) | 56.8% | |
| #24 | Grok 4.5 | 56.5% | |
| #25 | Claude Sonnet 5.5 (max effort) | 56.3% |
How it is scored
Arithmetic mean of the 3 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.
Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.