Math
FrontierMath Tier 4
The hardest FrontierMath set: exceptionally difficult research-level problems.
Published by Epoch AI under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| GPT-6.1 Sol (max) | 100.0% | #1 of 63 | |
| Claude Opus 5.5 (max) | 95.0% | #3 of 63 | |
| Gdm AI Co Mathematician | 75.6% | #10 of 63 | |
| Muse Spark 1.3 (max) | 46.3% | #19 of 63 | |
| Qwen3.8 Max (xhigh) | 46.3% | #19 of 63 | |
| Kimi K3 (max) | 39.0% | #22 of 63 | |
| Grok 4.6 (xhigh) | 31.7% | #26 of 63 | |
| GLM-5.2 (max) | 29.3% | #29 of 63 | |
| DeepSeek V4 Pro 0813 (max) | 26.8% | #32 of 63 |
New top scores
Each time a different model took the top spot, since the site began keeping this record.
- GPT-6.1 Sol took the top spot at 100.0% from GPT-6 Astra, which had led at 97.6%.
Every result
63 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | GPT-6.1 Sol (max) | 100.0% | |
| #2 | GPT-6 Astra (high) | 97.6% | |
| #3 | Claude Opus 5.5 (max) | 95.0% | |
| #4 | Claude Fable 5 (max) | 90.2% | |
| #5 | GPT-6 Sol (max) | 90.0% | |
| #6 | Claude Fable 5.1 (max) | 87.8% | |
| #7 | GPT-5.6 Sol (max) | 82.9% | |
| #8 | Claude Sonnet 5.5 (max) | 80.5% | |
| #9 | GPT-5.5 Pro (xhigh) | 78.0% | |
| #10 | Gdm AI Co Mathematician | 75.6% | |
| #11 | Claude Opus 5 (max) | 73.2% | |
| #12 | GPT-5.5 (xhigh) | 72.5% | |
| #13 | GPT-5.6 Terra (max) | 70.7% | |
| #14 | GPT-5.6 Luna (max) | 61.0% | |
| #15 | GPT-5.4 Pro (xhigh) | 58.5% | |
| #16 | Claude Opus 4.8 (max) | 56.1% | |
| #16 | GPT-6 Luna (max) | 56.1% | |
| #18 | GPT-5.4 (xhigh) | 49.0% | |
| #19 | Muse Spark 1.3 (max) | 46.3% | |
| #19 | Qwen3.8 Max (xhigh) | 46.3% | |
| #21 | GPT-5.2 Pro (xhigh) | 46.0% | |
| #22 | Kimi K3 (max) | 39.0% | |
| #23 | Gemini 3.7 Flash (high) | 36.6% | |
| #24 | Qwen3.7 Max | 34.1% | |
| #24 | Qwen3.8 Max 0902 (xhigh) | 34.1% |
How it is scored
Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.
Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.