All tests

Math

Mock AIME 2024–2025

Competition mathematics in the style of the AIME exam, from the OTIS olympiad program. Every answer is a whole number from 0 to 999.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score100.0%Claude Fable 5 (high), Anthropic
Models tested198from 15 makers
Median score67.2%Half the models score above it
Random guessing0.1%What guessing alone would score

Top 15

Each model at its best configuration, coloured by maker. The dashed line marks what random guessing would score.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Fable 5 (high)100.0%#1 of 198
OpenAIGPT-5.5 Pre-release (xhigh)100.0%#1 of 198
AlibabaQwen3.8 Max 0902 (xhigh)100.0%#1 of 198
xAIGrok 4.6 (xhigh)99.2%#14 of 198
Meta AIMuse Spark 1.3 (xhigh)99.2%#14 of 198
GoogleGemini 3.8 Flash (high)98.9%#16 of 198
DeepSeekDeepSeek V4 Pro 0813 (max)98.6%#19 of 198
Moonshot AIKimi K3 (max)97.2%#26 of 198
Z.aiGLM-5.3-Flash (max)93.9%#41 of 198
NVIDIAnemotron-3-ultra86.7%#62 of 198
MiniMaxMiniMax-M371.1%#91 of 198
ByteDanceByteDance-Seed/Seed-OSS-36B-Instruct67.5%#99 of 198

Every result

198 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Fable 5 (high)Anthropic100.0%
#1Claude Fable 5.1 (max)Anthropic100.0%
#1Claude Opus 5.5 (max)Anthropic100.0%
#1Claude Sonnet 5.5 (max)Anthropic100.0%
#1GPT-5.5 Pre-release (xhigh)OpenAI100.0%
#1GPT-5.5 Pro Pre-release (xhigh)OpenAI100.0%
#1GPT-5.6 Sol (max)OpenAI100.0%
#1GPT-6 Astra (max)OpenAI100.0%
#1GPT-6 Sol (max)OpenAI100.0%
#1GPT-6.1 Sol (max)OpenAI100.0%
#1Qwen3.8 Max 0902 (xhigh)Alibaba100.0%
#12GPT-5.6 Terra (max)OpenAI99.7%
#13Qwen3.8 Max (xhigh)Alibaba99.4%
#14Grok 4.6 (xhigh)xAI99.2%
#14Muse Spark 1.3 (xhigh)Meta AI99.2%
#16Claude Opus 5 (max)Anthropic98.9%
#16Gemini 3.8 Flash (high)Google98.9%
#16GPT-6 Luna (max)OpenAI98.9%
#19DeepSeek V4 Pro 0813 (max)DeepSeek98.6%
#20Claude Opus 4.8 (max)Anthropic98.3%
#20GPT-5.6 Luna (max)OpenAI98.3%
#22Grok 4.7 (xhigh)xAI98.1%
#23Claude Opus 4.7 (xhigh)Anthropic97.8%
#24GPT-5.4 (high)OpenAI97.8%
#24Grok 4.5 (high)xAI97.8%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.