All tests

Agents

DeepSWE

Original, long software engineering tasks from active open-source projects in TypeScript, Go, Python, JavaScript and Rust. Every model runs in the same minimal agent, mini-swe-agent.

Published by Datacurve DeepSWE, compiled by Epoch AI under CC BY 4.0. Compare it with other tests

RSS feed
Top score74.1%GPT-6 Astra (xhigh), OpenAI
Models tested26from 8 makers
Median score61.2%Half the models score above it
Scale0 to 100%The share of tasks done right

Top 15

Each model at its best configuration, coloured by maker.

The best model from each maker

Each maker’s highest-scoring model on this test, and its rank.

MakerBest modelScoreRank
OpenAIGPT-6 Astra (xhigh)74.1%#1 of 26
GoogleGemini 3.8 Flash (high)73.8%#2 of 26
AnthropicClaude Opus 5 (max)73.7%#3 of 26
Z.aiGLM-5.3 (max)69.0%#7 of 26
Moonshot AIKimi K3 (max)68.5%#8 of 26
xAIGrok 4.6 (medium)67.5%#9 of 26
AlibabaQwen3.8 Max (xhigh)57.5%#15 of 26
Meta AIMuse Spark 1.2 (xhigh)54.9%#16 of 26

Every result

26 models, each once at its best configuration, highest score first. Models with the same score share a rank.

RankModelMakerScore
#1GPT-6 Astra (xhigh)OpenAI74.1%
#2Gemini 3.8 Flash (high)Google73.8%
#3Claude Opus 5 (max)Anthropic73.7%
#4GPT-5.6 Sol (max)OpenAI72.7%
#5Claude Fable 5 (xhigh)Anthropic69.9%
#6GPT-5.6 Terra (max)OpenAI69.6%
#7GLM-5.3 (max)Z.ai69.0%
#8Kimi K3 (max)Moonshot AI68.5%
#9Grok 4.6 (medium)xAI67.5%
#10GPT-5.6 Luna (max)OpenAI67.2%
#11GPT-5.5 (xhigh)OpenAI67.0%
#12Gemini 3.7 Flash (medium)Google65.5%
#13GLM-5.3-Flash (max)Z.ai63.4%
#14Claude Opus 4.8 (max)Anthropic59.0%
#15Qwen3.8 Max (xhigh)Alibaba57.5%
#16Muse Spark 1.2 (xhigh)Meta AI54.9%
#17Claude Sonnet 5 (max)Anthropic53.9%
#18Grok 4.5 (high)xAI53.8%
#19Muse Spark 1.1Meta AI53.3%
#20GPT-5.4 (xhigh)OpenAI51.8%
#21Gemini 3.6 Flash (high)Google46.7%
#22GLM-5.2 (max)Z.ai43.8%
#23Gemini 3.5 Flash (medium)Google37.4%
#24Kimi K2.7 CodeMoonshot AI30.5%
#25Claude Sonnet 4.6 (high)Anthropic29.9%

How it is scored

Pass@1 averaged over repeated runs, with average API cost, output tokens and agent steps per task, as published by Datacurve and compiled by Epoch AI. A suffix such as _high names the reasoning effort. Standard errors are derived from the published 95% intervals.

Results from Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), as published. Model names link to their pages where the site can match them; makers come from the model’s page, or else from the publisher.