Overall
Capability index
One general-capability score that combines results from many benchmarks. Higher is better. The scale has no fixed maximum, so compare models rather than reading it as a percentage.
Published by Epoch AI under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Claude Opus 5.5 | 167.3 | #1 of 274 | |
| GPT-6 Astra | 166.4 | #2 of 274 | |
| Kimi K3 | 157.4 | #15 of 274 | |
| Gemini 3.7 Flash | 157.3 | #16 of 274 | |
| Muse Spark 1.3 | 156.8 | #19 of 274 | |
| Grok 4.6 | 156.4 | #21 of 274 | |
| Qwen 3.8 Max | 156.4 | #22 of 274 | |
| GLM-5.3 | 155.6 | #27 of 274 | |
| DeepSeek V4 Pro 0813 | 155.3 | #29 of 274 | |
| MiniMax M3 | 146.9 | #68 of 274 | |
| Nemotron 3 Ultra | 146.2 | #77 of 274 | |
| Mistral Medium 3.5 | 141.3 | #110 of 274 |
New top scores
Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.
- Claude Opus 5.5 took the top spot at 167.3 from GPT-6 Astra, which had led at 166.6.Within margin
Every result
274 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Claude Opus 5.5 | 167.3164.1 to 171.7 | |
| #2 | GPT-6 Astra | 166.4163.3 to 171.5 | |
| #3 | GPT-6.1 Sol | 166.1162.9 to 170.6 | |
| #4 | Claude Sonnet 5.5 | 165.0161.7 to 169.1 | |
| #5 | Claude Fable 5.1 | 164.7161.6 to 168.8 | |
| #6 | Claude Opus 5 | 162.8160.1 to 166.5 | |
| #7 | GPT-6 Sol | 162.7160.1 to 166.6 | |
| #8 | GPT-5.5 Pro | 162.1159.0 to 165.8 | |
| #9 | Claude Fable 5 | 162.1159.3 to 166.2 | |
| #10 | GPT-5.6 Sol | 161.7159.3 to 165.1 | |
| #11 | GPT-5.6 Terra | 159.6157.2 to 162.7 | |
| #12 | GPT-5.5 | 159.1156.7 to 162.3 | |
| #13 | GPT-5.4 Pro | 158.9156.6 to 161.8 | |
| #14 | Claude Opus 4.8 | 158.2156.0 to 160.8 | |
| #15 | Kimi K3 | 157.4155.2 to 160.0 | |
| #16 | Gemini 3.7 Flash | 157.3155.4 to 159.9 | |
| #17 | GPT-5.4 | 156.8155.0 to 158.9 | |
| #18 | GPT-5.3 Codex | 156.8152.8 to 160.7 | |
| #19 | Muse Spark 1.3 | 156.8154.8 to 159.3 | |
| #20 | Gemini 3.8 Flash | 156.7154.4 to 160.3 | |
| #21 | Grok 4.6 | 156.4154.5 to 159.0 | |
| #22 | Qwen 3.8 Max | 156.4154.3 to 158.6 | |
| #23 | GPT-5.6 Luna | 156.4154.0 to 158.7 | |
| #24 | GPT-6 Luna | 156.3153.7 to 158.8 | |
| #25 | Claude Opus 4.7 | 156.3154.4 to 158.4 |
How it is scored
Epoch Capabilities Index (ECI) as published by Epoch AI, with its published confidence interval. It is a statistical combination of many benchmarks, not a single test.
Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.