All tests

Language

LiveBench language

Language puzzles, connections and text understanding.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score90.7%Claude Fable 5 (max effort), Anthropic
Models tested63from 11 makers
Median score79.7%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Fable 5 (max effort)90.7%#1 of 63
OpenAIGPT-6.1 Sol (max)90.1%#2 of 63
GoogleGemini 3.8 Flash (high)87.8%#6 of 63
Moonshot AIKimi K385.5%#10 of 63
xAIGrok 4.683.7%#17 of 63
DeepSeekDeepSeek V4 Pro 081382.1%#24 of 63
Z.aiGLM-5.379.9%#29 of 63
AlibabaQwen3.7 Max79.7%#31 of 63
MiniMaxMiniMax-M376.8%#42 of 63
NVIDIANemotron 3 Ultra 550B A55B70.8%#59 of 63

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Fable 5 (max effort)Anthropic90.7%
#2GPT-6.1 Sol (max)OpenAI90.1%
#3Claude Fable 5.1 (max effort)Anthropic89.5%
#4GPT-6 Astra (max)OpenAI89.4%
#5Claude Opus 5 (max effort)Anthropic88.7%
#6Gemini 3.8 Flash (high)Google87.8%
#7GPT-5.6 Sol (max)OpenAI87.7%
#8GPT-5.5 (xhigh)OpenAI87.4%
#9Claude Opus 5.5 (max effort)Anthropic86.3%
#10Kimi K3Moonshot AI85.5%
#11Gemini 3.7 Flash (high)Google85.5%
#12Gemini 3.1 ProGoogle85.4%
#13GPT-6 Sol (max)OpenAI85.3%
#14Gemini 3.5 Flash (high)Google84.6%
#15Smaug AgenticNot reported84.4%
#16Gemini 3.6 Flash (high)Google83.9%
#17Grok 4.6xAI83.7%
#18Claude Sonnet 5.5 (xhigh effort)Anthropic83.4%
#19Claude Opus 4.6 (thinking auto high effort)Anthropic83.3%
#20GPT-5.6 Terra (max)OpenAI82.9%
#21Grok 4.5xAI82.8%
#22Muse Spark 1.3 (xhigh)Meta AI82.8%
#23GPT-5.4 (xhigh)OpenAI82.6%
#24DeepSeek V4 Pro 0813DeepSeek82.1%
#25Claude Opus 4.5 (20251101) (thinking 64k high effort)Anthropic81.3%

How it is scored

Arithmetic mean of the 3 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.