Agents
DeepSWE
Original, long software engineering tasks from active open-source projects in TypeScript, Go, Python, JavaScript and Rust. Every model runs in the same minimal agent, mini-swe-agent.
Published by Datacurve DeepSWE, compiled by Epoch AI under CC BY 4.0. Compare it with other tests
Top 15
Each model at its best configuration, coloured by maker.
The best model from each maker
Each maker’s highest-scoring model on this test, and its rank.
| Maker | Best model | Score | Rank |
|---|---|---|---|
| GPT-6 Astra (xhigh) | 74.1% | #1 of 26 | |
| Gemini 3.8 Flash (high) | 73.8% | #2 of 26 | |
| Claude Opus 5 (max) | 73.7% | #3 of 26 | |
| GLM-5.3 (max) | 69.0% | #7 of 26 | |
| Kimi K3 (max) | 68.5% | #8 of 26 | |
| Grok 4.6 (medium) | 67.5% | #9 of 26 | |
| Qwen3.8 Max (xhigh) | 57.5% | #15 of 26 | |
| Muse Spark 1.2 (xhigh) | 54.9% | #16 of 26 |
Every result
26 models, each once at its best configuration, highest score first. Models with the same score share a rank.
| Rank | Model | Maker | Score |
|---|---|---|---|
| #1 | GPT-6 Astra (xhigh) | 74.1% | |
| #2 | Gemini 3.8 Flash (high) | 73.8% | |
| #3 | Claude Opus 5 (max) | 73.7% | |
| #4 | GPT-5.6 Sol (max) | 72.7% | |
| #5 | Claude Fable 5 (xhigh) | 69.9% | |
| #6 | GPT-5.6 Terra (max) | 69.6% | |
| #7 | GLM-5.3 (max) | 69.0% | |
| #8 | Kimi K3 (max) | 68.5% | |
| #9 | Grok 4.6 (medium) | 67.5% | |
| #10 | GPT-5.6 Luna (max) | 67.2% | |
| #11 | GPT-5.5 (xhigh) | 67.0% | |
| #12 | Gemini 3.7 Flash (medium) | 65.5% | |
| #13 | GLM-5.3-Flash (max) | 63.4% | |
| #14 | Claude Opus 4.8 (max) | 59.0% | |
| #15 | Qwen3.8 Max (xhigh) | 57.5% | |
| #16 | Muse Spark 1.2 (xhigh) | 54.9% | |
| #17 | Claude Sonnet 5 (max) | 53.9% | |
| #18 | Grok 4.5 (high) | 53.8% | |
| #19 | Muse Spark 1.1 | 53.3% | |
| #20 | GPT-5.4 (xhigh) | 51.8% | |
| #21 | Gemini 3.6 Flash (high) | 46.7% | |
| #22 | GLM-5.2 (max) | 43.8% | |
| #23 | Gemini 3.5 Flash (medium) | 37.4% | |
| #24 | Kimi K2.7 Code | 30.5% | |
| #25 | Claude Sonnet 4.6 (high) | 29.9% |
How it is scored
Pass@1 averaged over repeated runs, with average API cost, output tokens and agent steps per task, as published by Datacurve and compiled by Epoch AI. A suffix such as _high names the reasoning effort. Standard errors are derived from the published 95% intervals.
Results from Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.