All tests

Agents

LiveBench agentic coding

JavaScript, TypeScript and Python agentic coding tasks.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score77.3%DeepSeek V4.1 Flash (max), DeepSeek
Models tested63from 11 makers
Median score53.8%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
DeepSeekDeepSeek V4.1 Flash (max)77.3%#1 of 63
AnthropicClaude Opus 5.5 (max effort)71.7%#2 of 63
AlibabaQwen3.8 Max64.7%#6 of 63
Moonshot AIKimi K362.2%#9 of 63
Z.aiGLM-5.360.9%#14 of 63
GoogleGemini 3.7 Flash (high)58.3%#18 of 63
OpenAIGPT-6 Astra (max)57.3%#20 of 63
xAIGrok 4.657.0%#21 of 63
MiniMaxMiniMax-M340.7%#58 of 63
NVIDIANemotron 3 Ultra 550B A55B38.7%#61 of 63

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1DeepSeek V4.1 Flash (max)DeepSeek77.3%
#2Claude Opus 5.5 (max effort)Anthropic71.7%
#3Claude Fable 5.1 (max effort)Anthropic66.1%
#4Claude Opus 5 (max effort)Anthropic65.2%
#5DeepSeek V4 Flash Vision ExpDeepSeek65.1%
#6Qwen3.8 MaxAlibaba64.7%
#6Smaug AgenticNot reported64.7%
#8Muse Spark 1.3 (xhigh)Meta AI64.1%
#9Claude Fable 5 (max effort)Anthropic62.2%
#9Kimi K3Moonshot AI62.2%
#11Qwen3.8 Flash NextAlibaba61.6%
#12Qwen3.8 27BAlibaba61.4%
#13Smaug FlashNot reported61.1%
#14GLM-5.3Z.ai60.9%
#15Smaug MiniNot reported60.8%
#16Claude Sonnet 5 (xhigh effort)Anthropic59.4%
#17Muse Spark 1.1 (xhigh)Meta AI58.5%
#18Gemini 3.7 Flash (high)Google58.3%
#19Muse Spark 1.2 (xhigh)Meta AI57.6%
#20GPT-6 Astra (max)OpenAI57.3%
#21Grok 4.6xAI57.0%
#22GLM-5.3-FlashZ.ai56.8%
#22GPT-6.1 Sol (xhigh)OpenAI56.8%
#24Grok 4.5xAI56.5%
#25Claude Sonnet 5.5 (max effort)Anthropic56.3%

How it is scored

Arithmetic mean of the 3 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.