Math
MATH Level 5
The hardest level of the MATH competition problem set. Leading models now score near the top, so it separates older and smaller models best.
Published by Epoch AI under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| GPT-5 (high) | 98.1% | #1 of 97 | |
| Claude Sonnet 4.5 (20250929) (32K) | 97.7% | #5 of 97 | |
| qwen3-max-2025-09-23 | 97.1% | #6 of 97 | |
| DeepSeek R1 0528 | 96.6% | #7 of 97 | |
| Gemini 2.5 Pro Preview 0506 | 95.9% | #10 of 97 | |
| Grok 3 Mini Beta (low) | 90.9% | #16 of 97 | |
| Mistral Medium 3 | 81.6% | #29 of 97 | |
| Llama 4 Maverick 17B 128E Instruct FP8 | 73.0% | #33 of 97 | |
| Phi-4 | 64.9% | #39 of 97 |
Every result
97 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | GPT-5 (high) | 98.1% | |
| #2 | GPT-5 Mini (high) | 97.8% | |
| #3 | o4-mini (high) | 97.8% | |
| #4 | o3 (high) | 97.8% | |
| #5 | Claude Sonnet 4.5 (20250929) (32K) | 97.7% | |
| #6 | qwen3-max-2025-09-23 | 97.1% | |
| #7 | DeepSeek R1 0528 | 96.6% | |
| #8 | o3-mini (high) | 96.5% | |
| #9 | Claude Haiku 4.5 (20251001) (32K) | 96.4% | |
| #10 | Gemini 2.5 Pro Preview 0506 | 95.9% | |
| #11 | Gemini 2.5 Pro Preview 0325 | 95.6% | |
| #12 | GPT-5 Nano (medium) | 95.2% | |
| #13 | o1 (high) | 94.7% | |
| #14 | DeepSeek-R1 | 93.0% | |
| #15 | Claude Sonnet 3.7 (64K) | 91.2% | |
| #16 | Grok 3 Mini Beta (low) | 90.9% | |
| #17 | DeepSeek R1 Distill Llama 70B | 89.9% | |
| #18 | o1-mini (high) | 89.2% | |
| #19 | Grok 3 Beta | 88.8% | |
| #20 | Grok 3 Mini Beta (high) | 88.1% | |
| #21 | OpenAI GPT-4.1 Mini | 87.3% | |
| #22 | DeepSeek R1 Distill Qwen 14B | 87.1% | |
| #23 | Claude Opus 4 (20250514) | 85.0% | |
| #24 | Claude Sonnet 4 (20250514) | 84.4% | |
| #25 | Gemini 2.0 Pro 0205 | 83.5% |
How it is scored
Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.
Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.