Agents
SWE-bench Verified, bash only
The 500 SWE-bench Verified issues from real Python projects, with every model in the same minimal agent that can only run shell commands (mini-SWE-agent).
Published by SWE-bench under CC BY-NC 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Claude Opus 4.5 (high) | 76.8% | #1 of 42 | |
| Gemini 3 Flash Preview (high) | 75.8% | #2 of 42 | |
| MiniMax-M2.5 (high) | 75.8% | #2 of 42 | |
| GLM-5 (high) | 72.8% | #7 of 42 | |
| GPT-5.2 (high) | 72.8% | #7 of 42 | |
| Kimi K2.5 (high) | 70.8% | #11 of 42 | |
| DeepSeek V3.2 (high) | 70.0% | #12 of 42 | |
| Qwen3-Coder 480B-A35B Instruct | 55.4% | #25 of 42 | |
| Devstral 2 | 53.8% | #28 of 42 | |
| Llama 4 Maverick Instruct | 21.0% | #39 of 42 |
Every result
42 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Claude Opus 4.5 (high) | 76.8% | |
| #2 | Gemini 3 Flash Preview (high) | 75.8% | |
| #2 | MiniMax-M2.5 (high) | 75.8% | |
| #4 | Claude Opus 4.6 | 75.6% | |
| #5 | Claude Opus 4.5 (20251101) (medium) | 74.4% | |
| #6 | Gemini 3 Pro Preview | 74.2% | |
| #7 | GLM-5 (high) | 72.8% | |
| #7 | GPT-5.2 (high) | 72.8% | |
| #7 | GPT-5.2 Codex | 72.8% | |
| #10 | Claude Sonnet 4.5 (20250929) (high) | 71.4% | |
| #11 | Kimi K2.5 (high) | 70.8% | |
| #12 | DeepSeek V3.2 (high) | 70.0% | |
| #13 | Claude Opus 4 (20250514) | 67.6% | |
| #14 | Claude Haiku 4.5 (20251001) (high) | 66.6% | |
| #15 | GPT-5.1 (2025-11-13) (medium) | 66.0% | |
| #15 | GPT-5.1 Codex (medium) | 66.0% | |
| #17 | GPT-5 (medium) | 65.0% | |
| #18 | Claude Sonnet 4 (20250514) | 64.9% | |
| #19 | Kimi K2 Thinking | 63.4% | |
| #20 | MiniMax-M2 | 61.0% | |
| #21 | DeepSeek V3.2 Reasoner | 60.0% | |
| #22 | GPT-5 Mini (medium) | 59.8% | |
| #23 | o3 (20250416) | 58.4% | |
| #24 | Devstral Small 2512 | 56.4% | |
| #25 | GLM-4.6 | 55.4% |
How it is scored
Share of issues resolved, with average API cost and model calls per issue, as published by the SWE-bench team. The latest run of each model and reasoning effort is kept; a suffix such as _high names the effort.
Results from SWE-bench (CC BY-NC 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.