All tests

Reasoning

LiveBench reasoning

Logical, spatial and multi-step reasoning tasks.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score92.7%GPT-6 Astra (max), OpenAI
Models tested63from 11 makers
Median score85.8%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6 Astra (max)92.7%#1 of 63
AnthropicClaude Opus 5.5 (max effort)92.2%#3 of 63
Moonshot AIKimi K390.7%#8 of 63
xAIGrok 4.690.5%#10 of 63
GoogleGemini 3.8 Flash (high)89.3%#16 of 63
AlibabaQwen3.8 Max88.2%#21 of 63
DeepSeekDeepSeek V4.1 Flash (max)86.7%#28 of 63
Z.aiGLM-5.385.8%#32 of 63
NVIDIANemotron 3 Ultra 550B A55B74.7%#57 of 63
MiniMaxMiniMax-M374.5%#58 of 63

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6 Astra (max)OpenAI92.7%
#2GPT-6.1 Sol (max)OpenAI92.6%
#3Claude Opus 5.5 (max effort)Anthropic92.2%
#4Claude Fable 5.1 (max effort)Anthropic91.7%
#5GPT-5.6 Sol (max)OpenAI91.7%
#6Claude Sonnet 5.5 (max effort)Anthropic91.6%
#7Claude Opus 5 (max effort)Anthropic91.2%
#8Kimi K3Moonshot AI90.7%
#9GPT-5.6 Terra (max)OpenAI90.6%
#10Grok 4.6xAI90.5%
#11Smaug AgenticNot reported90.3%
#12Muse Spark 1.2 (xhigh)Meta AI90.0%
#13Claude Fable 5 (max effort)Anthropic89.7%
#13GPT-5.5 (xhigh)OpenAI89.7%
#13Muse Spark 1.3 (xhigh)Meta AI89.7%
#16Gemini 3.8 Flash (high)Google89.3%
#17Claude Opus 4.8 (max effort)Anthropic89.2%
#18Claude Sonnet 5 (xhigh effort)Anthropic88.7%
#19Claude Opus 4.6 (thinking auto high effort)Anthropic88.7%
#20GPT-6 Sol (max)OpenAI88.7%
#21Qwen3.8 MaxAlibaba88.2%
#22GPT-5.4 (xhigh)OpenAI88.1%
#23Gemini 3.7 Flash (high)Google87.8%
#24Muse Spark 1.1 (xhigh)Meta AI87.7%
#25Qwen3.8 Flash NextAlibaba87.4%

How it is scored

Arithmetic mean of the 4 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.