All tests

Agents

SWE-bench Verified, bash only

The 500 SWE-bench Verified issues from real Python projects, with every model in the same minimal agent that can only run shell commands (mini-SWE-agent).

Published by SWE-bench under CC BY-NC 4.0. Compare it with other tests

RSS feed
Top score76.8%Claude Opus 4.5 (high), Anthropic
Models tested42from 11 makers
Median score59.9%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
AnthropicClaude Opus 4.5 (high)76.8%#1 of 42
GoogleGemini 3 Flash Preview (high)75.8%#2 of 42
MiniMaxMiniMax-M2.5 (high)75.8%#2 of 42
Z.aiGLM-5 (high)72.8%#7 of 42
OpenAIGPT-5.2 (high)72.8%#7 of 42
Moonshot AIKimi K2.5 (high)70.8%#11 of 42
DeepSeekDeepSeek V3.2 (high)70.0%#12 of 42
AlibabaQwen3-Coder 480B-A35B Instruct55.4%#25 of 42
Mistral AIDevstral 253.8%#28 of 42
MetaLlama 4 Maverick Instruct21.0%#39 of 42

Every result

42 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1Claude Opus 4.5 (high)Anthropic76.8%
#2Gemini 3 Flash Preview (high)Google75.8%
#2MiniMax-M2.5 (high)MiniMax75.8%
#4Claude Opus 4.6Anthropic75.6%
#5Claude Opus 4.5 (20251101) (medium)Anthropic74.4%
#6Gemini 3 Pro PreviewGoogle74.2%
#7GLM-5 (high)Z.ai72.8%
#7GPT-5.2 (high)OpenAI72.8%
#7GPT-5.2 CodexOpenAI72.8%
#10Claude Sonnet 4.5 (20250929) (high)Anthropic71.4%
#11Kimi K2.5 (high)Moonshot AI70.8%
#12DeepSeek V3.2 (high)DeepSeek70.0%
#13Claude Opus 4 (20250514)Anthropic67.6%
#14Claude Haiku 4.5 (20251001) (high)Anthropic66.6%
#15GPT-5.1 (2025-11-13) (medium)OpenAI66.0%
#15GPT-5.1 Codex (medium)OpenAI66.0%
#17GPT-5 (medium)OpenAI65.0%
#18Claude Sonnet 4 (20250514)Anthropic64.9%
#19Kimi K2 ThinkingMoonshot AI63.4%
#20MiniMax-M2MiniMax61.0%
#21DeepSeek V3.2 Reasonerdeepseek60.0%
#22GPT-5 Mini (medium)OpenAI59.8%
#23o3 (20250416)OpenAI58.4%
#24Devstral Small 2512mistral56.4%
#25GLM-4.6Z.ai55.4%

How it is scored

Share of issues resolved, with average API cost and model calls per issue, as published by the SWE-bench team. The latest run of each model and reasoning effort is kept; a suffix such as _high names the effort.

Results from SWE-bench (CC BY-NC 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.