Math
Mock AIME 2024–2025
Competition mathematics in the style of the AIME exam, from the OTIS olympiad program. Every answer is a whole number from 0 to 999.
Published by Epoch AI under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The dashed line marks what random guessing would score.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Claude Fable 5 (high) | 100.0% | #1 of 198 | |
| GPT-5.5 Pre-release (xhigh) | 100.0% | #1 of 198 | |
| Qwen3.8 Max 0902 (xhigh) | 100.0% | #1 of 198 | |
| Grok 4.6 (xhigh) | 99.2% | #14 of 198 | |
| Muse Spark 1.3 (xhigh) | 99.2% | #14 of 198 | |
| Gemini 3.8 Flash (high) | 98.9% | #16 of 198 | |
| DeepSeek V4 Pro 0813 (max) | 98.6% | #19 of 198 | |
| Kimi K3 (max) | 97.2% | #26 of 198 | |
| GLM-5.3-Flash (max) | 93.9% | #41 of 198 | |
| nemotron-3-ultra | 86.7% | #62 of 198 | |
| MiniMax-M3 | 71.1% | #91 of 198 | |
| ByteDance-Seed/Seed-OSS-36B-Instruct | 67.5% | #99 of 198 |
Every result
198 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Claude Fable 5 (high) | 100.0% | |
| #1 | Claude Fable 5.1 (max) | 100.0% | |
| #1 | Claude Opus 5.5 (max) | 100.0% | |
| #1 | Claude Sonnet 5.5 (max) | 100.0% | |
| #1 | GPT-5.5 Pre-release (xhigh) | 100.0% | |
| #1 | GPT-5.5 Pro Pre-release (xhigh) | 100.0% | |
| #1 | GPT-5.6 Sol (max) | 100.0% | |
| #1 | GPT-6 Astra (max) | 100.0% | |
| #1 | GPT-6 Sol (max) | 100.0% | |
| #1 | GPT-6.1 Sol (max) | 100.0% | |
| #1 | Qwen3.8 Max 0902 (xhigh) | 100.0% | |
| #12 | GPT-5.6 Terra (max) | 99.7% | |
| #13 | Qwen3.8 Max (xhigh) | 99.4% | |
| #14 | Grok 4.6 (xhigh) | 99.2% | |
| #14 | Muse Spark 1.3 (xhigh) | 99.2% | |
| #16 | Claude Opus 5 (max) | 98.9% | |
| #16 | Gemini 3.8 Flash (high) | 98.9% | |
| #16 | GPT-6 Luna (max) | 98.9% | |
| #19 | DeepSeek V4 Pro 0813 (max) | 98.6% | |
| #20 | Claude Opus 4.8 (max) | 98.3% | |
| #20 | GPT-5.6 Luna (max) | 98.3% | |
| #22 | Grok 4.7 (xhigh) | 98.1% | |
| #23 | Claude Opus 4.7 (xhigh) | 97.8% | |
| #24 | GPT-5.4 (high) | 97.8% | |
| #24 | Grok 4.5 (high) | 97.8% |
How it is scored
Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.
Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.