All tests

Language

LiveBench instructions

Following constraints in rewriting, summaries and generation.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score81.4%Gemini 3.8 Flash (high), Google
Models tested63from 11 makers
Median score69.3%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
GoogleGemini 3.8 Flash (high)81.4%#1 of 63
AlibabaQwen3.8 Flash Next77.1%#5 of 63
AnthropicClaude Fable 5 (max effort)75.8%#6 of 63
OpenAIGPT-6 Astra (max)75.6%#8 of 63
xAIGrok 4.7 (xhigh)75.3%#10 of 63
NVIDIANemotron 3 Ultra 550B A55B73.5%#16 of 63
Moonshot AIKimi K371.4%#23 of 63
DeepSeekDeepSeek V4 Flash Vision Exp71.0%#25 of 63
Z.aiGLM-5.369.3%#32 of 63
MiniMaxMiniMax-M357.5%#59 of 63

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Gemini 3.8 Flash (high)Google81.4%
#2Gemini 3.7 Flash (high)Google79.9%
#3Gemini 3.1 ProGoogle79.1%
#4Muse Spark 1.3 (xhigh)Meta AI78.0%
#5Qwen3.8 Flash NextAlibaba77.1%
#6Claude Fable 5 (max effort)Anthropic75.8%
#7Gemini 3.5 Flash (high)Google75.6%
#8GPT-6 Astra (max)OpenAI75.6%
#9Gemini 3.6 Flash (high)Google75.4%
#10Grok 4.7 (xhigh)xAI75.3%
#11Muse Spark 1.2 (xhigh)Meta AI74.3%
#12GPT-6.1 Sol (max)OpenAI74.2%
#13Qwen3.8 MaxAlibaba74.1%
#14Qwen3.7 MaxAlibaba74.0%
#15Smaug MiniNot reported73.9%
#16Nemotron 3 Ultra 550B A55BNVIDIA73.5%
#17Claude Fable 5.1 (max effort)Anthropic73.0%
#18Qwen3.8 27BAlibaba72.7%
#19Claude Opus 4.8 (max effort)Anthropic72.0%
#20Grok 4.6xAI71.9%
#21GPT-5.6 Sol (max)OpenAI71.8%
#22Grok 4.5xAI71.5%
#23Kimi K3Moonshot AI71.4%
#24Smaug AgenticNot reported71.0%
#25DeepSeek V4 Flash Vision ExpDeepSeek71.0%

How it is scored

Arithmetic mean of the 4 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.