All tests

Math

FrontierMath Tier 4

The hardest FrontierMath set: exceptionally difficult research-level problems.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score100.0%GPT-6.1 Sol (max), OpenAI
Models tested63from 10 makers
Median score26.8%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6.1 Sol (max)100.0%#1 of 63
AnthropicClaude Opus 5.5 (max)95.0%#3 of 63
Google DeepMindGdm AI Co Mathematician75.6%#10 of 63
Meta AIMuse Spark 1.3 (max)46.3%#19 of 63
AlibabaQwen3.8 Max (xhigh)46.3%#19 of 63
Moonshot AIKimi K3 (max)39.0%#22 of 63
xAIGrok 4.6 (xhigh)31.7%#26 of 63
Z.aiGLM-5.2 (max)29.3%#29 of 63
DeepSeekDeepSeek V4 Pro 0813 (max)26.8%#32 of 63

New top scores

Each time a different model took the top spot, since the site began keeping this record.

  1. GPT-6.1 Sol took the top spot at 100.0% from GPT-6 Astra, which had led at 97.6%.

Every result

63 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6.1 Sol (max)OpenAI100.0%
#2GPT-6 Astra (high)OpenAI97.6%
#3Claude Opus 5.5 (max)Anthropic95.0%
#4Claude Fable 5 (max)Anthropic90.2%
#5GPT-6 Sol (max)OpenAI90.0%
#6Claude Fable 5.1 (max)Anthropic87.8%
#7GPT-5.6 Sol (max)OpenAI82.9%
#8Claude Sonnet 5.5 (max)Anthropic80.5%
#9GPT-5.5 Pro (xhigh)OpenAI78.0%
#10Gdm AI Co MathematicianGoogle DeepMind75.6%
#11Claude Opus 5 (max)Anthropic73.2%
#12GPT-5.5 (xhigh)OpenAI72.5%
#13GPT-5.6 Terra (max)OpenAI70.7%
#14GPT-5.6 Luna (max)OpenAI61.0%
#15GPT-5.4 Pro (xhigh)OpenAI58.5%
#16Claude Opus 4.8 (max)Anthropic56.1%
#16GPT-6 Luna (max)OpenAI56.1%
#18GPT-5.4 (xhigh)OpenAI49.0%
#19Muse Spark 1.3 (max)Meta AI46.3%
#19Qwen3.8 Max (xhigh)Alibaba46.3%
#21GPT-5.2 Pro (xhigh)OpenAI46.0%
#22Kimi K3 (max)Moonshot AI39.0%
#23Gemini 3.7 Flash (high)Google36.6%
#24Qwen3.7 MaxAlibaba34.1%
#24Qwen3.8 Max 0902 (xhigh)Alibaba34.1%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.