All tests

Reasoning

GPQA Diamond

Graduate-level multiple-choice questions in biology, chemistry and physics, written by experts to be hard to answer with a web search. Random guessing scores 25%.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score95.8%GPT-6 Astra (max), OpenAI
Models tested222from 15 makers
Median score67.0%Half the models score above it
Random guessing25.0%What guessing alone would score

Top 15

Each model at its best configuration, coloured by maker. The dashed line marks what random guessing would score.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6 Astra (max)95.8%#1 of 222
AnthropicClaude Sonnet 5.5 (max)95.6%#2 of 222
GoogleGemini 3.8 Flash (high)95.4%#3 of 222
xAIGrok 4.6 (high)94.0%#11 of 222
Moonshot AIKimi K3 (max)93.1%#19 of 222
AlibabaQwen3.8 Max (xhigh)92.7%#22 of 222
Z.aiGLM-5.2 (max)91.9%#25 of 222
DeepSeekDeepSeek V4 Pro 0813 (max)91.7%#26 of 222
MiniMaxMiniMax-M390.9%#31 of 222
Meta AIMuse Spark89.8%#44 of 222
NVIDIAnemotron-3-ultra85.3%#65 of 222
ByteDanceByteDance-Seed/Seed-OSS-36B-Instruct71.5%#103 of 222

Every result

222 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6 Astra (max)OpenAI95.8%
#2Claude Sonnet 5.5 (max)Anthropic95.6%
#3Gemini 3.8 Flash (high)Google95.4%
#3GPT-6.1 Sol (max)OpenAI95.4%
#5Gemini 3.7 Flash (high)Google94.8%
#6GPT-5.4 Pro (xhigh)OpenAI94.6%
#7Gemini 3.1 Pro (high)Google94.4%
#8GPT-6 Sol (max)OpenAI94.3%
#9Gemini 3.6 Flash (high)Google94.1%
#10Gemini 3.1 Pro PreviewGoogle94.1%
#11GPT-5.5 Pre-release (xhigh)OpenAI94.0%
#11Grok 4.6 (high)xAI94.0%
#13GPT-5.5 Pro Pre-release (xhigh)OpenAI93.9%
#14Claude Opus 5 (max)Anthropic93.9%
#15GPT-5.6 Sol (max)OpenAI93.5%
#16Grok 4.5 (high)xAI93.4%
#17GPT-5.6 Terra (max)OpenAI93.3%
#18GPT-5.4 (xhigh)OpenAI93.3%
#19Kimi K3 (max)Moonshot AI93.1%
#20Gemini 3.5 Flash (high)Google92.8%
#21Grok 4.7 (xhigh)xAI92.7%
#22Qwen3.8 Max (xhigh)Alibaba92.7%
#23Gemini 3 Pro PreviewGoogle92.6%
#24Qwen3.8 Max 0902 (xhigh)Alibaba92.3%
#25GLM-5.2 (max)Z.ai91.9%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.