All tests

Math

LiveBench mathematics

Competition mathematics, symbolic work and problem solving.

Published by LiveBench under Apache-2.0. Compare it with other tests

RSS feed
Top score97.1%Claude Opus 5.5 (max effort), Anthropic
Models tested63from 11 makers
Median score89.8%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Opus 5.5 (max effort)97.1%#1 of 63
OpenAIGPT-6.1 Sol (max)96.8%#3 of 63
xAIGrok 4.7 (xhigh)95.7%#12 of 63
DeepSeekDeepSeek V4 Pro 081395.1%#13 of 63
GoogleGemini 3.7 Flash (high)93.5%#17 of 63
AlibabaQwen3.8 Max91.3%#24 of 63
Z.aiGLM-5.289.8%#32 of 63
NVIDIANemotron 3 Ultra 550B A55B88.7%#37 of 63
Moonshot AIKimi K384.4%#50 of 63
MiniMaxMiniMax-M377.0%#62 of 63

New top scores

Each time a different model took the top spot, since the site began keeping this record.

  1. Claude Opus 5.5 took the top spot at 97.1% from Claude Fable 5.1, which had led at 97.0%.

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Opus 5.5 (max effort)Anthropic97.1%
#2Claude Fable 5.1 (max effort)Anthropic97.0%
#3GPT-6.1 Sol (max)OpenAI96.8%
#4GPT-6 Astra (max)OpenAI96.8%
#5Claude Sonnet 5.5 (xhigh effort)Anthropic96.7%
#6GPT-6 Sol (max)OpenAI96.4%
#7GPT-5.6 Sol (max)OpenAI96.2%
#8Claude Fable 5 (max effort)Anthropic96.0%
#9Muse Spark 1.3 (xhigh)Meta AI96.0%
#10GPT-5.5 (xhigh)OpenAI95.9%
#11Claude Opus 5 (max effort)Anthropic95.7%
#12Grok 4.7 (xhigh)xAI95.7%
#13DeepSeek V4 Pro 0813DeepSeek95.1%
#14GPT-5.6 Terra (max)OpenAI94.9%
#15Claude Opus 4.8 (max effort)Anthropic94.3%
#16GPT-5.4 (xhigh)OpenAI94.2%
#17Gemini 3.7 Flash (high)Google93.5%
#18DeepSeek V4.1 Flash (max)DeepSeek93.3%
#19GPT-5.2 (2025 12 11 high)OpenAI93.2%
#20Claude Sonnet 5 (xhigh effort)Anthropic92.9%
#21Claude Opus 4.7 (xhigh effort)Anthropic92.8%
#22Grok 4.6xAI92.6%
#23Gemini 3.8 Flash (high)Google91.6%
#24Qwen3.8 MaxAlibaba91.3%
#25Muse Spark 1.2 (xhigh)Meta AI91.2%

How it is scored

Arithmetic mean of the 4 published task scores in this category. Only complete results are ranked. Exact model and effort configurations are preserved.

Results from LiveBench (Apache-2.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.