All tests

Math

FrontierMath Tiers 1–3

Original, unpublished mathematics problems from advanced undergraduate to research level, written by mathematicians and checked automatically.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score93.7%GPT-6 Astra (max), OpenAI
Models tested81from 10 makers
Median score55.8%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6 Astra (max)93.7%#1 of 81
AnthropicClaude Opus 5.5 (max)91.2%#3 of 81
AlibabaQwen3.8 Max (xhigh)74.7%#18 of 81
Meta AIMuse Spark 1.3 (xhigh)74.4%#19 of 81
Moonshot AIKimi K3 (max)72.2%#21 of 81
GoogleGemini 3.7 Flash (high)71.6%#22 of 81
Z.aiGLM-5.3 (max)68.8%#24 of 81
xAIGrok 4.6 (xhigh)66.0%#27 of 81
DeepSeekDeepSeek V4 Pro 0813 (max)64.6%#31 of 81

Every result

81 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6 Astra (max)OpenAI93.7%
#1GPT-6.1 Sol (max)OpenAI93.7%
#3Claude Opus 5.5 (max)Anthropic91.2%
#4Claude Fable 5.1 (max)Anthropic90.2%
#5GPT-6 Sol (max)OpenAI89.8%
#6GPT-5.6 Sol (max)OpenAI89.1%
#7Claude Sonnet 5.5 (max)Anthropic88.8%
#8GPT-5.5 Pro (xhigh)OpenAI87.7%
#9Claude Fable 5 (max)Anthropic87.0%
#10GPT-5.6 Terra (max)OpenAI86.0%
#11Claude Opus 5 (max)Anthropic85.6%
#12GPT-5.5 (xhigh)OpenAI85.3%
#13GPT-5.4 Pro (xhigh)OpenAI82.5%
#14GPT-5.6 Luna (max)OpenAI82.1%
#15Claude Opus 4.8 (max)Anthropic80.0%
#16GPT-6 Luna (max)OpenAI79.0%
#17GPT-5.4 (xhigh)OpenAI78.6%
#18Qwen3.8 Max (xhigh)Alibaba74.7%
#19Muse Spark 1.3 (xhigh)Meta AI74.4%
#20GPT-5.2 Pro (xhigh)OpenAI74.0%
#21Kimi K3 (max)Moonshot AI72.2%
#22Gemini 3.7 Flash (high)Google71.6%
#23Claude Opus 4.7 (max)Anthropic70.2%
#24GLM-5.3 (max)Z.ai68.8%
#25Gemini 3.8 Flash (high)Google68.4%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.