All tests

Coding

LiveBench coding

Code generation and completion on the published LiveBench tasks.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score91.4%Claude Sonnet 5.5 (max effort), Anthropic
Models tested63from 11 makers
Median score77.9%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Sonnet 5.5 (max effort)91.4%#1 of 63
OpenAIGPT-5.6 Sol (max)83.9%#5 of 63
Moonshot AIKimi K381.5%#13 of 63
DeepSeekDeepSeek V4.1 Flash (max)80.0%#19 of 63
Z.aiGLM-5.279.7%#20 of 63
GoogleGemini 3.7 Flash (high)78.9%#26 of 63
AlibabaQwen3.6 Plus78.2%#29 of 63
xAIGrok 4.7 (xhigh)77.2%#35 of 63
NVIDIANemotron 3 Ultra 550B A55B70.7%#56 of 63
MiniMaxMiniMax-M368.2%#61 of 63

New top scores

Each time a different model took the top spot, since the site began keeping this record.

  1. Claude Sonnet 5.5 took the top spot at 91.4% from Claude Opus 5.5, which had led at 89.3%.
  2. Claude Opus 5.5 took the top spot at 89.3% from Claude Fable 5.1, which had led at 86.4%.

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Sonnet 5.5 (max effort)Anthropic91.4%
#2Claude Opus 5.5 (max effort)Anthropic89.3%
#3Claude Fable 5.1 (max effort)Anthropic86.4%
#4Claude Fable 5 (max effort)Anthropic86.0%
#5GPT-5.6 Sol (max)OpenAI83.9%
#6GPT-5.2 CodexOpenAI83.6%
#7GPT-5.6 Luna (max)OpenAI82.9%
#8Smaug AgenticNot reported82.5%
#9GPT-5.5 (xhigh)OpenAI82.2%
#10Claude Opus 4.7 (xhigh effort)Anthropic82.1%
#11Claude Opus 4.8 (max effort)Anthropic81.8%
#12GPT-6 Sol (max)OpenAI81.8%
#13Claude Opus 5 (max effort)Anthropic81.5%
#13Kimi K3Moonshot AI81.5%
#15Muse Spark 1.3 (xhigh)Meta AI81.1%
#16GPT-6.1 Sol (xhigh)OpenAI80.7%
#17Claude Sonnet 5 (xhigh effort)Anthropic80.7%
#18GPT-6 Astra (max)OpenAI80.4%
#19DeepSeek V4.1 Flash (max)DeepSeek80.0%
#20Claude Opus 4.5 (20251101) (thinking 64k high effort)Anthropic79.7%
#20GLM-5.2Z.ai79.7%
#22Claude Sonnet 4.6 (thinking auto medium effort)Anthropic79.3%
#23GLM-5.3Z.ai79.0%
#23GLM-5.3-FlashZ.ai79.0%
#23GPT-6 Luna (max)OpenAI79.0%

How it is scored

Arithmetic mean of the 2 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.