All tests

Knowledge

SimpleQA Verified

Short factual questions with a single correct answer, testing whether a model recalls facts accurately.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score75.6%GPT-6 Astra (max), OpenAI
Models tested77from 10 makers
Median score44.1%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6 Astra (max)75.6%#1 of 77
GoogleGemini 3.1 Pro (high)73.5%#3 of 77
AnthropicClaude Opus 5.5 (max)72.2%#4 of 77
Meta AIMuse Spark 1.2 (xhigh)60.3%#15 of 77
xAIGrok 4.7 (xhigh)56.0%#18 of 77
AlibabaQwen3.7 Max55.8%#19 of 77
DeepSeekDeepSeek V4 Pro 0813 (max)52.9%#21 of 77
Moonshot AIKimi K3 (max)50.6%#24 of 77
Z.aiGLM-5.3 (max)41.0%#43 of 77

Every result

77 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6 Astra (max)OpenAI75.6%
#2GPT-6.1 Sol (max)OpenAI73.9%
#3Gemini 3.1 Pro (high)Google73.5%
#4Claude Opus 5.5 (max)Anthropic72.2%
#5Claude Fable 5.1 (max)Anthropic70.8%
#6Claude Fable 5 (xhigh)Anthropic70.7%
#7Gemini 3.8 Flash (high)Google69.7%
#7GPT-5.6 Sol (max)OpenAI69.7%
#9Gemini 3.7 Flash (high)Google69.2%
#10Gemini 3 Flash Preview (high)Google66.8%
#11Gemini 3.5 Flash (high)Google66.2%
#11Gemini 3.6 Flash (high)Google66.2%
#13GPT-5.5 (xhigh)OpenAI63.0%
#14GPT-6 Sol (max)OpenAI60.7%
#15Muse Spark 1.2 (xhigh)Meta AI60.3%
#16Claude Opus 5 (max)Anthropic59.9%
#17Muse Spark 1.1Meta AI57.8%
#18Grok 4.7 (xhigh)xAI56.0%
#19Qwen3.7 MaxAlibaba55.8%
#20Claude Opus 4.8 (max)Anthropic53.0%
#21DeepSeek V4 Pro 0813 (max)DeepSeek52.9%
#22Qwen3.6 Max PreviewAlibaba52.0%
#23Claude Opus 4.7 (xhigh)Anthropic51.7%
#24Kimi K3 (max)Moonshot AI50.6%
#25GPT-5 (high)OpenAI50.1%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.