All tests

Overall

LiveBench overall

Broad language-model performance across the current LiveBench release.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score83.4%Claude Fable 5.1 (max effort), Anthropic
Models tested63from 11 makers
Median score75.8%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Fable 5.1 (max effort)83.4%#1 of 63
OpenAIGPT-6 Astra (max)82.2%#4 of 63
DeepSeekDeepSeek V4.1 Flash (max)81.1%#7 of 63
Moonshot AIKimi K379.2%#13 of 63
GoogleGemini 3.7 Flash (high)78.8%#14 of 63
AlibabaQwen3.8 Max78.5%#15 of 63
xAIGrok 4.678.0%#16 of 63
Z.aiGLM-5.376.1%#30 of 63
NVIDIANemotron 3 Ultra 550B A55B67.4%#57 of 63
MiniMaxMiniMax-M367.3%#58 of 63

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Fable 5.1 (max effort)Anthropic83.4%
#2Claude Opus 5.5 (max effort)Anthropic83.2%
#3Claude Fable 5 (max effort)Anthropic83.0%
#4GPT-6 Astra (max)OpenAI82.2%
#5GPT-6.1 Sol (max)OpenAI81.6%
#6Muse Spark 1.3 (xhigh)Meta AI81.6%
#7DeepSeek V4.1 Flash (max)DeepSeek81.1%
#8GPT-5.6 Sol (max)OpenAI81.0%
#9GPT-5.5 (xhigh)OpenAI80.2%
#10Claude Opus 5 (max effort)Anthropic80.1%
#11Smaug AgenticNot reported79.5%
#12GPT-6 Sol (max)OpenAI79.3%
#13Kimi K3Moonshot AI79.2%
#14Gemini 3.7 Flash (high)Google78.8%
#15Qwen3.8 MaxAlibaba78.5%
#16Grok 4.6xAI78.0%
#17GPT-5.4 (xhigh)OpenAI78.0%
#18Muse Spark 1.2 (xhigh)Meta AI78.0%
#19GPT-5.6 Terra (max)OpenAI77.9%
#20Claude Sonnet 5.5 (xhigh effort)Anthropic77.8%
#21DeepSeek V4 Pro 0813DeepSeek77.4%
#21Smaug FlashNot reported77.4%
#23Grok 4.7 (xhigh)xAI77.4%
#24Gemini 3.1 ProGoogle77.0%
#25Smaug MiniNot reported76.9%

How it is scored

Equal-weight mean of all 7 category means, matching the published aggregation method for complete results. This is LiveBench performance, not a universal best-model score.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.