All tests

Math

MATH Level 5

The hardest level of the MATH competition problem set. Leading models now score near the top, so it separates older and smaller models best.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score98.1%GPT-5 (high), OpenAI
Models tested97from 10 makers
Median score52.6%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-5 (high)98.1%#1 of 97
AnthropicClaude Sonnet 4.5 (20250929) (32K)97.7%#5 of 97
Alibabaqwen3-max-2025-09-2397.1%#6 of 97
DeepSeekDeepSeek R1 052896.6%#7 of 97
GoogleGemini 2.5 Pro Preview 050695.9%#10 of 97
xAIGrok 3 Mini Beta (low)90.9%#16 of 97
Mistral AIMistral Medium 381.6%#29 of 97
MetaLlama 4 Maverick 17B 128E Instruct FP873.0%#33 of 97
MicrosoftPhi-464.9%#39 of 97

Every result

97 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-5 (high)OpenAI98.1%
#2GPT-5 Mini (high)OpenAI97.8%
#3o4-mini (high)OpenAI97.8%
#4o3 (high)OpenAI97.8%
#5Claude Sonnet 4.5 (20250929) (32K)Anthropic97.7%
#6qwen3-max-2025-09-23Alibaba97.1%
#7DeepSeek R1 0528DeepSeek96.6%
#8o3-mini (high)OpenAI96.5%
#9Claude Haiku 4.5 (20251001) (32K)Anthropic96.4%
#10Gemini 2.5 Pro Preview 0506Google95.9%
#11Gemini 2.5 Pro Preview 0325Google95.6%
#12GPT-5 Nano (medium)OpenAI95.2%
#13o1 (high)OpenAI94.7%
#14DeepSeek-R1DeepSeek93.0%
#15Claude Sonnet 3.7 (64K)Anthropic91.2%
#16Grok 3 Mini Beta (low)xAI90.9%
#17DeepSeek R1 Distill Llama 70BDeepSeek89.9%
#18o1-mini (high)OpenAI89.2%
#19Grok 3 BetaxAI88.8%
#20Grok 3 Mini Beta (high)xAI88.1%
#21OpenAI GPT-4.1 MiniOpenAI87.3%
#22DeepSeek R1 Distill Qwen 14BDeepSeek87.1%
#23Claude Opus 4 (20250514)Anthropic85.0%
#24Claude Sonnet 4 (20250514)Anthropic84.4%
#25Gemini 2.0 Pro 0205Google83.5%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.