All tests

Reasoning

Chess puzzles

Chess puzzles given to the model as text, scored on finding the correct moves.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score72.0%GPT-6 Astra (max), OpenAI
Models tested141from 15 makers
Median score17.0%Half the models score above it
Random guessing5.0%What guessing alone would score

Top 15

Each model at its best configuration, coloured by maker. The dashed line marks what random guessing would score.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6 Astra (max)72.0%#1 of 141
GoogleGemini 3.8 Flash (high)61.0%#4 of 141
AnthropicClaude Fable 5.1 (max)47.0%#13 of 141
DeepSeekDeepSeek V4 Pro 0813 (max)47.0%#13 of 141
xAIGrok 4.6 (high)40.0%#20 of 141
AlibabaQwen3.8 Max 0902 (xhigh)40.0%#20 of 141
Moonshot AIKimi K3 (max)39.0%#24 of 141
Meta AIMuse Spark 1.3 (max)38.0%#25 of 141
Z.aiGLM-5.2 (max)21.0%#55 of 141
MiniMaxMiniMax-M314.0%#73 of 141
ByteDanceByteDance-Seed/Seed-OSS-36B-Instruct13.0%#76 of 141
NVIDIAnemotron-3-ultra12.0%#80 of 141

Every result

141 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6 Astra (max)OpenAI72.0%
#2GPT-5.5 Pro Pre-release (xhigh)OpenAI64.0%
#2GPT-5.6 Sol (promax)OpenAI64.0%
#4Gemini 3.8 Flash (high)Google61.0%
#4GPT-6.1 Sol (max)OpenAI61.0%
#6GPT-5.4 Pro (xhigh)OpenAI58.6%
#7Gemini 3.1 Pro PreviewGoogle55.0%
#8GPT-5.5 Pre-release (xhigh)OpenAI54.0%
#8GPT-5.6 Terra (max)OpenAI54.0%
#10Gemini 3.5 Flash (high)Google50.0%
#11Gemini 3.1 Pro (high)Google49.0%
#11GPT-5.2 (xhigh)OpenAI49.0%
#13Claude Fable 5.1 (max)Anthropic47.0%
#13DeepSeek V4 Pro 0813 (max)DeepSeek47.0%
#13Gemini 3.7 Flash (high)Google47.0%
#16GPT-5.4 (xhigh)OpenAI44.0%
#17Gemini 3.6 Flash (low)Google43.0%
#18Claude Opus 5 (max)Anthropic42.0%
#19Claude Fable 5 (high)Anthropic41.0%
#20Gemini 3 Flash Preview (high)Google40.0%
#20GPT-5.6 Luna (max)OpenAI40.0%
#20Grok 4.6 (high)xAI40.0%
#20Qwen3.8 Max 0902 (xhigh)Alibaba40.0%
#24Kimi K3 (max)Moonshot AI39.0%
#25Grok 4.7 (xhigh)xAI38.0%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.