All tests

Data analysis

LiveBench data analysis

Working with tables and structured information.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score83.0%GPT-6 Astra (max), OpenAI
Models tested63from 11 makers
Median score76.2%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6 Astra (max)83.0%#1 of 63
AnthropicClaude Fable 5 (max effort)80.5%#5 of 63
DeepSeekDeepSeek V4 Flash Vision Exp79.5%#11 of 63
Moonshot AIKimi K378.7%#19 of 63
GoogleGemini 3.1 Pro78.5%#21 of 63
AlibabaQwen3.8 Max78.4%#22 of 63
xAIGrok 4.7 (xhigh)76.9%#28 of 63
Z.aiGLM-5.3-Flash76.4%#31 of 63
MiniMaxMiniMax-M376.2%#32 of 63
NVIDIANemotron 3 Ultra 550B A55B54.5%#61 of 63

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6 Astra (max)OpenAI83.0%
#2GPT-6.1 Sol (max)OpenAI82.7%
#3GPT-5.5 (xhigh)OpenAI81.6%
#4GPT-6 Sol (max)OpenAI81.2%
#5Claude Fable 5 (max effort)Anthropic80.5%
#6Claude Opus 5.5 (max effort)Anthropic80.3%
#7Claude Fable 5.1 (max effort)Anthropic80.3%
#8Smaug AgenticNot reported79.9%
#9GPT-5.6 Sol (max)OpenAI79.8%
#10Muse Spark 1.3 (xhigh)Meta AI79.6%
#11DeepSeek V4 Flash Vision ExpDeepSeek79.5%
#12DeepSeek V4 Flash 0731DeepSeek79.3%
#13GPT-5.4 (xhigh)OpenAI79.3%
#13GPT-5.6 Terra (max)OpenAI79.3%
#15DeepSeek V4.1 Flash (max)DeepSeek79.3%
#16DeepSeek V4 Pro 0813DeepSeek79.2%
#17Smaug FlashNot reported79.1%
#18Smaug MiniNot reported78.8%
#19Kimi K3Moonshot AI78.7%
#20Claude Sonnet 5.5 (xhigh effort)Anthropic78.6%
#21Gemini 3.1 ProGoogle78.5%
#22Qwen3.8 MaxAlibaba78.4%
#23Claude Opus 4.7 (xhigh effort)Anthropic78.3%
#24GPT-5.2 CodexOpenAI78.2%
#25GPT-5.2 (2025 12 11 high)OpenAI78.2%

How it is scored

Arithmetic mean of the 3 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.