All tests

Coding

SWE-bench Verified

Real issues from open-source Python projects. The model must change the code so the project’s tests pass. Engineers confirmed each task is solvable.

Published by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score83.5%Claude Opus 4.7 (max), Anthropic
Models tested32from 7 makers
Median score74.0%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Opus 4.7 (max)83.5%#1 of 32
OpenAIGPT-5.5 Pre-release (xhigh)80.6%#2 of 32
GoogleGemini 3.5 Flash (high)79.3%#3 of 32
Z.aiGLM-5.2 (max)78.7%#5 of 32
DeepSeekDeepSeek V4 Pro (max)77.6%#6 of 32
AlibabaQwen3.7 Max77.3%#7 of 32
Moonshot AIKimi K2.676.7%#9 of 32

Every result

32 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Opus 4.7 (max)Anthropic83.5%
#2GPT-5.5 Pre-release (xhigh)OpenAI80.6%
#3Gemini 3.5 Flash (high)Google79.3%
#4Claude Opus 4.6Anthropic78.7%
#5GLM-5.2 (max)Z.ai78.7%
#6DeepSeek V4 Pro (max)DeepSeek77.6%
#7Qwen3.7 MaxAlibaba77.3%
#8GPT-5.4 (high)OpenAI76.9%
#9Claude Opus 4.5 (20251101)Anthropic76.7%
#9Kimi K2.6Moonshot AI76.7%
#9Qwen3.6 Max PreviewAlibaba76.7%
#12Gemini 3.1 Pro Preview Custom ToolsGoogle75.6%
#13Gemini 3 Flash PreviewGoogle75.4%
#14Claude Sonnet 4.6Anthropic75.2%
#15GPT-5.3 Codex (high)OpenAI74.8%
#16GLM-5.1Z.ai74.2%
#17GPT-5.2 (high)OpenAI73.8%
#17Kimi K2.5Moonshot AI73.8%
#19GPT-5 (high)OpenAI73.5%
#20Claude Opus 4.1 (20250805)Anthropic73.3%
#21Gemini 3 Pro PreviewGoogle72.9%
#22GLM-5Z.ai72.1%
#23Claude Sonnet 4.5 (20250929)Anthropic71.3%
#24Claude Opus 4 (20250514)Anthropic70.7%
#25GPT-5.1 (2025-11-13) (high)OpenAI68.0%

How it is scored

Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.

Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.