Coding agents and the models inside them.
Public test results for AI coding agents: which agent and model pairs solve the most tasks, how much the agent changes a model’s score, and which models lead on coding tests.
At a glance
What the published results say about agents and the models inside them.
- Best pair
NexAU-AHE with
GPT-5.5 (84.7%) solves the most Terminal-Bench 2.0 tasks, ahead of Capy with
GPT-5.5 (83.2%).
- Maker’s own agent
Among agents built by the model’s maker, Codex CLI with
GPT-5.5 (82.3%) scores highest.
- Agent effect
Across 21 models run by at least three agents, the best agent typically beats the weakest by 13.3 points. For
Claude Opus 4.6 the gap is 21.8 points: ForgeCode 79.8%, Claude Code 58.0%.
- Cost per task
On DeepSWE,
GPT-6 Astra (xhigh) scores highest (74.1%) at $6.52 a task, while
Gemini 3.8 Flash (medium) comes within five points (71.0%) for $1.97.
- Task length
Claude Mythos Preview completes half of the software tasks that take a skilled person 17h 25m, the longest in METR’s measurements.
- Most tested agent
Terminus 2 has Terminal-Bench 2.0 results with 33 models, more than any other agent.
- Newest models
On LiveBench agentic coding,
DeepSeek V4.1 Flash (max) and
Claude Opus 5.5 (max effort) lead, followed by
Claude Fable 5.1 (max effort) and
Claude Opus 5 (max effort).
Leaderboard
Terminal-Bench 2.0 results for 141 agent and model pairs. Hard tasks carried out in a computer terminal, from building and debugging software to data and system work, each checked automatically.
Agent and model pairs
Share of the 89 Terminal-Bench 2.0 tasks solved. Higher is better. Every result pairs one agent with one model.
- NexAU-AHEchina-qijizhifeng
withGPT-5.584.7%plus or minus 2.1
- CapyCapy
withGPT-5.583.2%plus or minus 2.1
- Codex CLIOpenAI
withGPT-5.582.3%plus or minus 2.2
- CodexOpenAI
withGPT-5.582.0%plus or minus 2.2
- ForgeCodeForgeCode
withGPT-5.4 (2026-03-05)81.8%plus or minus 2.0
- TongAgentsBigai
withGemini 3.1 Pro Preview80.2%plus or minus 2.6
- WOZCODEWOZCODE
withClaude Opus 4.780.2%plus or minus 2.1
- ForgeCodeForgeCode
withClaude Opus 4.679.8%plus or minus 1.6
- SageAgentOpenSage
withGPT-5.3 Codex78.4%plus or minus 2.2
- ForgeCodeForgeCode
withGemini 3.1 Pro Preview78.4%plus or minus 1.8
- DroidFactory
withGPT-5.3 Codex77.3%plus or minus 2.2
- Meta-HarnessStanford IRIS
withClaude Opus 4.676.4%plus or minus 2.4
- CodeBrain-1.5Feeling AI
withGPT-5.3 Codex75.8%plus or minus 2.0
- Codeliakousw
withGPT-5.3 Codex75.7%plus or minus 2.2
- CapyCapy
withClaude Opus 4.675.3%plus or minus 2.4
Bars start at zero; the thin line marks one standard error either side. Terminal-Bench 2.0 (Apache-2.0), compiled by Epoch AI (CC BY 4.0).
How much the agent matters
The same model can solve far more tasks in one agent than in another.
Same model, different agents
Each line runs from the weakest to the strongest agent’s score for one model; every dot is an agent. Open a model to see which agent scored what.
GPT-5.55 agents, 18.7 points apart66.1% to 84.7%
- NexAU-AHEchina-qijizhifeng84.7%
- CapyCapy83.2%
- Codex CLIOpenAI82.3%
- CodexOpenAI82.0%
- clnkrclnkr66.1%
Gemini 3.1 Pro Preview4 agents, 20.9 points apart59.4% to 80.2%
- TongAgentsBigai80.2%
- ForgeCodeForgeCode78.4%
- Terminus-KIRAKRAFTON AI74.8%
- Gemini CLIGoogle59.4%
Claude Opus 4.611 agents, 21.8 points apart58.0% to 79.8%
- ForgeCodeForgeCode79.8%
- Meta-HarnessStanford IRIS76.4%
- CapyCapy75.3%
- Terminus-KIRAKRAFTON AI74.7%
- MAYA-V2ADYA72.1%
- TongAgentsBigai71.9%
- DroidFactory69.9%
- CruxRoam66.9%
- MuxCoder66.5%
- Terminus 2Terminal-Bench62.9%
- Claude CodeAnthropic58.0%
GPT-5.3 Codex9 agents, 13.7 points apart64.7% to 78.4%
- SageAgentOpenSage78.4%
- DroidFactory77.3%
- CodeBrain-1.5Feeling AI75.8%
- Codeliakousw75.7%
- Simple CodexOpenAI75.1%
- MuxCoder74.6%
- spoox-o-mTUM71.5%
- IndusAGI Coding AgentVarun Israni (SoloVpx)69.1%
- Terminus 2Terminal-Bench64.7%
Gemini 3 Pro Preview7 agents, 13.5 points apart56.0% to 69.4%
- AnteAntigma Labs69.4%
- SageAgentOpenSage65.2%
- CodeBrain-1.5Feeling AI62.3%
- II-AgentIntelligent Internet61.8%
- DroidFactory61.1%
- Terminus 2Terminal-Bench56.9%
- Letta CodeLetta56.0%
Gemini 3 Flash Preview3 agents, 13.3 points apart51.0% to 64.3%
- Junie CLIJetBrains64.3%
- Terminus 2Terminal-Bench51.7%
- Gemini CLIGoogle51.0%
Claude Opus 4.5 (20251101)8 agents, 11.4 points apart51.7% to 63.1%
- DroidFactory63.1%
- Letta CodeLetta59.1%
- MuxCoder58.4%
- Terminus 2Terminal-Bench57.8%
- GooseBlock54.3%
- Claude CodeAnthropic52.1%
- OpenHandsOpenHands51.9%
- OpenCodeAnomaly Innovations51.7%
GPT-5.1 Codex mini3 agents, 18.5 points apart43.1% to 61.6%
- hookeleDmitry Barakhov61.6%
- CruxRoam43.1%
- Codex CLIOpenAI43.1%
GPT-5.1 Codex4 agents, 20.9 points apart36.9% to 57.8%
- Codex CLIOpenAI57.8%
- CruxRoam57.8%
- Letta CodeLetta53.5%
- Terminus 2Terminal-Bench36.9%
GPT-5 (medium)4 agents, 15.7 points apart33.9% to 49.6%
- Codex CLIOpenAI49.6%
- OpenHandsOpenHands43.8%
- Terminus 2Terminal-Bench35.2%
- Mini-SWE-AgentPrinceton33.9%
Models with results from at least three agents. Terminal-Bench 2.0 (Apache-2.0), compiled by Epoch AI (CC BY 4.0).
Agents
All 40 agents with Terminal-Bench 2.0 results, with who built them and their best result.
- NexAU-AHEchina-qijizhifengModels run1Best result84.7%
withGPT-5.5Latest resultApr 23, 2026
- CapyCapyModels run2Best result83.2%
withGPT-5.5Latest resultApr 23, 2026
- Codex CLIOpenAIModels run9Best result82.3%
withGPT-5.5Latest resultApr 23, 2026
- CodexOpenAIModels run1Best result82.0%
withGPT-5.5Latest resultApr 23, 2026
- ForgeCodeForgeCodeModels run3Best result81.8%
withGPT-5.4 (2026-03-05)Latest resultMar 12, 2026
- TongAgentsBigaiModels run2Best result80.2%
withGemini 3.1 Pro PreviewLatest resultFeb 19, 2026
- WOZCODEWOZCODEModels run1Best result80.2%
withClaude Opus 4.7Latest resultApr 16, 2026
- SageAgentOpenSageModels run2Best result78.4%
withGPT-5.3 CodexLatest resultFeb 5, 2026
- DroidFactoryModels run5Best result77.3%
withGPT-5.3 CodexLatest resultFeb 5, 2026
- Meta-HarnessStanford IRISModels run1Best result76.4%
withClaude Opus 4.6Latest resultFeb 5, 2026
- CodeBrain-1.5Feeling AIModels run2Best result75.8%
withGPT-5.3 CodexLatest resultFeb 5, 2026
- CodeliakouswModels run1Best result75.7%
withGPT-5.3 CodexLatest resultFeb 5, 2026
Sorted by each agent’s best result.
Cost per task
What models cost and how much work they do per task, as the tests measured it.
DeepSWE: score against cost per task
What one task cost on average, against the share of tasks solved. Up and to the left is better value; the line joins the models no cheaper run beats.
OpenAI
Anthropic
Google
xAI
Z.ai
Meta AI
Moonshot AI
Alibaba
- Best score at each cost
- Above-median score, below-median cost
Use the arrow keys to move between models. Press Enter to open the selected model. A table with the same data follows the chart.
| Model | Score | Cost per task | Output tokens per task | Agent steps per task |
|---|---|---|---|---|
| GPT-6 Astra (xhigh) | 74.1% | $6.52 | 29.6K | 28.8 |
| Gemini 3.8 Flash (high) | 73.8% | $2.36 | 143K | 166.3 |
| Claude Opus 5 (max) | 73.7% | $11.84 | 118K | 99 |
| GPT-6 Astra (high) | 73.2% | $5.72 | 26.5K | 27.4 |
| GPT-6 Astra (max) | 73.2% | $12.37 | 61.1K | 28.5 |
| Claude Opus 5 (xhigh) | 73.2% | $9.07 | 91.7K | 88.7 |
| Claude Opus 5 (high) | 72.8% | $6.08 | 64.2K | 72.9 |
| GPT-6 Astra (medium) | 72.8% | $4.38 | 20.4K | 26 |
| GPT-5.6 Sol (max) | 72.7% | $8.39 | 60K | 61.3 |
| Gemini 3.8 Flash (medium) | 71.0% | $1.97 | 125K | 147.3 |
| GPT-5.6 Sol (xhigh) | 70.7% | $4.70 | 40.7K | 44 |
| Claude Fable 5 (xhigh) | 69.9% | $13.41 | 80.4K | 68.4 |
Measured and published by Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0). Cost is the API cost of the tokens a run used, at list prices. Tasks and agents differ between tests, so compare costs within one test only.
How long a task models can finish
METR’s time horizons: longer tasks, measured in human time, that models complete.
METR times software tasks with skilled people, then finds the length of task each model completes half the time. The ranges are wide: for Claude Mythos Preview the 95% interval runs from 8h 29m to 55h 4m, and the task length it completes four times out of five is 3h 6m. Only results from the latest method, METR-Horizon-v1.1, are shown. Data: METR time horizons (compiled by Epoch AI, CC BY 4.0).
Models on coding tests
How the models themselves compare when every model runs in the same setup, including releases newer than the agent results.
Arena WebDev
People describe a web page or app; two anonymous models each build it, and people vote for the better working result.
Each model shows its best tested configuration. A test runs every model the same way, so these charts compare models, not agents. Results published by Arena (CC BY 4.0).
All 123 resultsCoding score and price
Which models give the most coding ability for their list price.
Arena WebDev against price
Up and to the left is better value: a higher score for a lower list price. The line joins the models no cheaper model beats.
Anthropic
Alibaba
Google
OpenAI
xAI
Z.ai
DeepSeek
Moonshot AI
Xiaomi
MiniMax
Mistral AI
Meta
Thinking Machines
Tencent
Inception
StepFun
Upstage
- Best score at each price
- Above-median score, below-median price
Use the arrow keys to move between models. Press Enter to open the selected model. A table with the same data follows the chart.
Price is the model’s list price per 1M tokens, blending three input tokens with one output token. It is not the cost of a task: an agent can use many tokens on one task. Scores from Arena (CC BY 4.0) and models.dev (MIT).
Frequently asked questions
Answered from the data on this page.
What is an AI coding agent?
A coding agent is the program that turns an AI model into something that can work on code: it gives the model tools to read and edit files, run commands and tests, and keep going until the task is done. Codex CLI, Claude Code, Gemini CLI and OpenHands are examples. The same model can do better or worse depending on the agent around it.
Which coding agent scores highest?
On Terminal-Bench 2.0, NexAU-AHE running GPT-5.5 solves 84.7% of the tasks, the highest result. The best agent from a model maker is Codex CLI with GPT-5.5, at 82.3%.
Does the agent matter as much as the model?
It can. For the 21 models that at least three agents ran, the gap between the best and weakest agent has a median of 13.3 points and reaches 21.8 points for Claude Opus 4.6, from 58.0% with Claude Code to 79.8% with ForgeCode.
Which agent gets the most from GPT-5.5?
NexAU-AHE 84.7%, Capy 83.2%, Codex CLI 82.3%, Codex 82.0%, clnkr 66.1%. NexAU-AHE and Capy are within their standard errors of each other, so treat them as tied.
How much does an AI coding agent cost per task?
It depends on the model, its reasoning effort and the task. DeepSWE, CursorBench, SWE-bench Verified, bash only publish the average API cost per task. On DeepSWE, GPT-6 Astra (xhigh) scores highest at $6.52 a task, while Gemini 3.8 Flash (medium) comes within five points for $1.97. Costs use list prices and differ between tests, so compare them within one test.
How long a task can an AI agent finish?
METR measures the length of software task, in the time a skilled person needs, that a model completes half the time. Claude Mythos Preview leads its measurements at 17h 25m. The uncertainty is large, and the task length a model completes four times out of five is much shorter.
Why are the newest models missing from the agent results?
Agent and model pairs appear here only after the Terminal-Bench 2.0 leaderboard lists them, which can take weeks after a model is released. For newer models, the model tests on this page, such as Arena WebDev, DeepSWE, CursorBench and LiveBench agentic coding, are more current.
Which model is best at agentic coding?
On LiveBench agentic coding, DeepSeek V4.1 Flash (max) scores highest at 77.3%, ahead of Claude Opus 5.5 (max effort) at 71.7%. Other agentic tests on this page rank models differently, so check the test closest to your work.
Which AI is best at building websites and web apps?
On Arena WebDev, where people compare two web apps built from the same request without knowing which model made them, Claude Opus 5.5 (max) scores highest at 1815, ahead of GPT-6 Astra (max) (1788). The scores come from blind votes, so they measure which result people prefer.
Where does this data come from?
Agent results come from the Terminal-Bench 2.0 leaderboard (Apache-2.0, compiled by Epoch AI) and SWE-bench (CC BY-NC 4.0). Model tests come from Arena (CC BY 4.0), LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), SWE-bench, and DeepSWE by Datacurve and CursorBench by Cursor as compiled by Epoch AI; task lengths come from METR, also via Epoch AI. Prices are from models.dev (MIT). Artificials does not run these tests; it matches results to models and calculates the rankings, gaps and charts. Everything refreshes automatically.
Sources and method
Artificials does not run these tests. Rankings, the agent gaps, the matching of results to models and the charts are calculated by Artificials from published results. Some sources allow only non-commercial use; Artificials is a free, non-commercial site.
- Agent results: Terminal-Bench 2.0 (Apache-2.0), as compiled by Epoch AI (CC BY 4.0).
- Agent results: SWE-bench Verified (CC BY-NC 4.0).
- Task lengths: METR time horizons (METR-Horizon-v1.1), as compiled by Epoch AI (CC BY 4.0).
- Model tests: LiveBench (Apache-2.0).
- Model tests: Epoch AI (CC BY 4.0).
- Model tests: Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0).
- Model tests: Cursor CursorBench, compiled by Epoch AI (CC BY 4.0).
- Model tests: Arena (CC BY 4.0).
- Model tests: SWE-bench (CC BY-NC 4.0).
- Prices: models.dev (MIT), the maker’s own list price or else the middle price across hosts.
DeepSWE, CursorBench and METR publish their results without a licence of their own; Epoch AI republishes them under CC BY 4.0, and they are credited to both. Agent names are merged when a source spells one agent several ways, and an entry listed twice keeps its latest run. Whether an agent comes from the model’s maker is worked out from the builder the source names. SWE-bench agent results include only entries that tried each issue once with one model.