Coding
SWE-bench Verified
Real issues from open-source Python projects. The model must change the code so the project’s tests pass. Engineers confirmed each task is solvable.
Published by Epoch AI under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| Claude Opus 4.7 (max) | 83.5% | #1 of 32 | |
| GPT-5.5 Pre-release (xhigh) | 80.6% | #2 of 32 | |
| Gemini 3.5 Flash (high) | 79.3% | #3 of 32 | |
| GLM-5.2 (max) | 78.7% | #5 of 32 | |
| DeepSeek V4 Pro (max) | 77.6% | #6 of 32 | |
| Qwen3.7 Max | 77.3% | #7 of 32 | |
| Kimi K2.6 | 76.7% | #9 of 32 |
Every result
32 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | Claude Opus 4.7 (max) | 83.5% | |
| #2 | GPT-5.5 Pre-release (xhigh) | 80.6% | |
| #3 | Gemini 3.5 Flash (high) | 79.3% | |
| #4 | Claude Opus 4.6 | 78.7% | |
| #5 | GLM-5.2 (max) | 78.7% | |
| #6 | DeepSeek V4 Pro (max) | 77.6% | |
| #7 | Qwen3.7 Max | 77.3% | |
| #8 | GPT-5.4 (high) | 76.9% | |
| #9 | Claude Opus 4.5 (20251101) | 76.7% | |
| #9 | Kimi K2.6 | 76.7% | |
| #9 | Qwen3.6 Max Preview | 76.7% | |
| #12 | Gemini 3.1 Pro Preview Custom Tools | 75.6% | |
| #13 | Gemini 3 Flash Preview | 75.4% | |
| #14 | Claude Sonnet 4.6 | 75.2% | |
| #15 | GPT-5.3 Codex (high) | 74.8% | |
| #16 | GLM-5.1 | 74.2% | |
| #17 | GPT-5.2 (high) | 73.8% | |
| #17 | Kimi K2.5 | 73.8% | |
| #19 | GPT-5 (high) | 73.5% | |
| #20 | Claude Opus 4.1 (20250805) | 73.3% | |
| #21 | Gemini 3 Pro Preview | 72.9% | |
| #22 | GLM-5 | 72.1% | |
| #23 | Claude Sonnet 4.5 (20250929) | 71.3% | |
| #24 | Claude Opus 4 (20250514) | 70.7% | |
| #25 | GPT-5.1 (2025-11-13) (high) | 68.0% |
How it is scored
Best published score for each exact model configuration, as evaluated by Epoch AI. A suffix such as _high names the reasoning effort used. Standard errors are shown where the source reports them.
Results from Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.