All tests

Overall

Capability index

One general-capability score that combines results from many benchmarks. Higher is better. The scale has no fixed maximum, so compare models rather than reading it as a percentage.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score167.3Claude Opus 5.5, Anthropic4 other models within its margin
Models tested274from 15 makers
Median score135.9Half the models score above it
ScaleOpen-endedCompare models, not the number itself

Top 15

Each model at its best configuration, coloured by maker. The publisher gives each score a range; a model whose score falls inside the leader’s range is not clearly behind it.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Opus 5.5167.3#1 of 274
OpenAIGPT-6 Astra166.4#2 of 274
Moonshot AIKimi K3157.4#15 of 274
GoogleGemini 3.7 Flash157.3#16 of 274
Meta AIMuse Spark 1.3156.8#19 of 274
xAIGrok 4.6156.4#21 of 274
AlibabaQwen 3.8 Max156.4#22 of 274
Z.aiGLM-5.3155.6#27 of 274
DeepSeekDeepSeek V4 Pro 0813155.3#29 of 274
MiniMaxMiniMax M3146.9#68 of 274
NVIDIANemotron 3 Ultra146.2#77 of 274
Mistral AIMistral Medium 3.5141.3#110 of 274

New top scores

Each time a different model took the top spot, since the site began keeping this record. “Within margin” means the next model’s score was inside the range the publisher gives for the new leader’s score.

  1. Claude Opus 5.5 took the top spot at 167.3 from GPT-6 Astra, which had led at 166.6.Within margin

Every result

274 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Opus 5.5Anthropic167.3164.1 to 171.7
#2GPT-6 AstraOpenAI166.4163.3 to 171.5
#3GPT-6.1 SolOpenAI166.1162.9 to 170.6
#4Claude Sonnet 5.5Anthropic165.0161.7 to 169.1
#5Claude Fable 5.1Anthropic164.7161.6 to 168.8
#6Claude Opus 5Anthropic162.8160.1 to 166.5
#7GPT-6 SolOpenAI162.7160.1 to 166.6
#8GPT-5.5 ProOpenAI162.1159.0 to 165.8
#9Claude Fable 5Anthropic162.1159.3 to 166.2
#10GPT-5.6 SolOpenAI161.7159.3 to 165.1
#11GPT-5.6 TerraOpenAI159.6157.2 to 162.7
#12GPT-5.5OpenAI159.1156.7 to 162.3
#13GPT-5.4 ProOpenAI158.9156.6 to 161.8
#14Claude Opus 4.8Anthropic158.2156.0 to 160.8
#15Kimi K3Moonshot AI157.4155.2 to 160.0
#16Gemini 3.7 FlashGoogle157.3155.4 to 159.9
#17GPT-5.4OpenAI156.8155.0 to 158.9
#18GPT-5.3 CodexOpenAI156.8152.8 to 160.7
#19Muse Spark 1.3Meta AI156.8154.8 to 159.3
#20Gemini 3.8 FlashGoogle156.7154.4 to 160.3
#21Grok 4.6xAI156.4154.5 to 159.0
#22Qwen 3.8 MaxAlibaba156.4154.3 to 158.6
#23GPT-5.6 LunaOpenAI156.4154.0 to 158.7
#24GPT-6 LunaOpenAI156.3153.7 to 158.8
#25Claude Opus 4.7Anthropic156.3154.4 to 158.4

How it is scored

Epoch Capabilities Index (ECI) as published by Epoch AI, with its published confidence interval. It is a statistical combination of many benchmarks, not a single test.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.