CODING AGENTS

Coding agents and the models inside them.

Public test results for AI coding agents: which agent and model pairs solve the most tasks, how much the agent changes a model’s score, and which models lead on coding tests.

At a glance

What the published results say about agents and the models inside them.

Best pair

NexAU-AHE with GPT-5.5 (84.7%) solves the most Terminal-Bench 2.0 tasks, ahead of Capy with GPT-5.5 (83.2%).

Maker’s own agent

Among agents built by the model’s maker, Codex CLI with GPT-5.5 (82.3%) scores highest.

Agent effect

Across 21 models run by at least three agents, the best agent typically beats the weakest by 13.3 points. For Claude Opus 4.6 the gap is 21.8 points: ForgeCode 79.8%, Claude Code 58.0%.

Cost per task

On DeepSWE, GPT-6 Astra (xhigh) scores highest (74.1%) at $6.52 a task, while Gemini 3.8 Flash (medium) comes within five points (71.0%) for $1.97.

Task length

Claude Mythos Preview completes half of the software tasks that take a skilled person 17h 25m, the longest in METR’s measurements.

Most tested agent

Terminus 2 has Terminal-Bench 2.0 results with 33 models, more than any other agent.

Newest models

On LiveBench agentic coding, DeepSeek V4.1 Flash (max) and Claude Opus 5.5 (max effort) lead, followed by Claude Fable 5.1 (max effort) and Claude Opus 5 (max effort).

Leaderboard

Terminal-Bench 2.0 results for 141 agent and model pairs. Hard tasks carried out in a computer terminal, from building and debugging software to data and system work, each checked automatically.

Agent and model pairs

Share of the 89 Terminal-Bench 2.0 tasks solved. Higher is better. Every result pairs one agent with one model.

  1. NexAU-AHEchina-qijizhifengwithGPT-5.584.7%plus or minus 2.1
  2. CapyCapywithGPT-5.583.2%plus or minus 2.1
  3. Codex CLIOpenAIwithGPT-5.582.3%plus or minus 2.2
  4. CodexOpenAIwithGPT-5.582.0%plus or minus 2.2
  5. ForgeCodeForgeCodewithGPT-5.4 (2026-03-05)81.8%plus or minus 2.0
  6. TongAgentsBigaiwithGemini 3.1 Pro Preview80.2%plus or minus 2.6
  7. WOZCODEWOZCODEwithClaude Opus 4.780.2%plus or minus 2.1
  8. ForgeCodeForgeCodewithClaude Opus 4.679.8%plus or minus 1.6
  9. SageAgentOpenSagewithGPT-5.3 Codex78.4%plus or minus 2.2
  10. ForgeCodeForgeCodewithGemini 3.1 Pro Preview78.4%plus or minus 1.8
  11. DroidFactorywithGPT-5.3 Codex77.3%plus or minus 2.2
  12. Meta-HarnessStanford IRISwithClaude Opus 4.676.4%plus or minus 2.4
  13. CodeBrain-1.5Feeling AIwithGPT-5.3 Codex75.8%plus or minus 2.0
  14. CodeliakouswwithGPT-5.3 Codex75.7%plus or minus 2.2
  15. CapyCapywithClaude Opus 4.675.3%plus or minus 2.4

Bars start at zero; the thin line marks one standard error either side. Terminal-Bench 2.0 (Apache-2.0), compiled by Epoch AI (CC BY 4.0).

How much the agent matters

The same model can solve far more tasks in one agent than in another.

Same model, different agents

Each line runs from the weakest to the strongest agent’s score for one model; every dot is an agent. Open a model to see which agent scored what.

  • GPT-5.55 agents, 18.7 points apart66.1% to 84.7%
    1. NexAU-AHEchina-qijizhifeng84.7%
    2. CapyCapy83.2%
    3. Codex CLIOpenAI82.3%
    4. CodexOpenAI82.0%
    5. clnkrclnkr66.1%
    GPT-5.5 profile
  • Gemini 3.1 Pro Preview4 agents, 20.9 points apart59.4% to 80.2%
    1. TongAgentsBigai80.2%
    2. ForgeCodeForgeCode78.4%
    3. Terminus-KIRAKRAFTON AI74.8%
    4. Gemini CLIGoogle59.4%
    Gemini 3.1 Pro Preview profile
  • Claude Opus 4.611 agents, 21.8 points apart58.0% to 79.8%
    1. ForgeCodeForgeCode79.8%
    2. Meta-HarnessStanford IRIS76.4%
    3. CapyCapy75.3%
    4. Terminus-KIRAKRAFTON AI74.7%
    5. MAYA-V2ADYA72.1%
    6. TongAgentsBigai71.9%
    7. DroidFactory69.9%
    8. CruxRoam66.9%
    9. MuxCoder66.5%
    10. Terminus 2Terminal-Bench62.9%
    11. Claude CodeAnthropic58.0%
    Claude Opus 4.6 profile
  • GPT-5.3 Codex9 agents, 13.7 points apart64.7% to 78.4%
    1. SageAgentOpenSage78.4%
    2. DroidFactory77.3%
    3. CodeBrain-1.5Feeling AI75.8%
    4. Codeliakousw75.7%
    5. Simple CodexOpenAI75.1%
    6. MuxCoder74.6%
    7. spoox-o-mTUM71.5%
    8. IndusAGI Coding AgentVarun Israni (SoloVpx)69.1%
    9. Terminus 2Terminal-Bench64.7%
    GPT-5.3 Codex profile
  • Gemini 3 Pro Preview7 agents, 13.5 points apart56.0% to 69.4%
    1. AnteAntigma Labs69.4%
    2. SageAgentOpenSage65.2%
    3. CodeBrain-1.5Feeling AI62.3%
    4. II-AgentIntelligent Internet61.8%
    5. DroidFactory61.1%
    6. Terminus 2Terminal-Bench56.9%
    7. Letta CodeLetta56.0%
    Gemini 3 Pro Preview profile
  • Gemini 3 Flash Preview3 agents, 13.3 points apart51.0% to 64.3%
    1. Junie CLIJetBrains64.3%
    2. Terminus 2Terminal-Bench51.7%
    3. Gemini CLIGoogle51.0%
    Gemini 3 Flash Preview profile
  • Claude Opus 4.5 (20251101)8 agents, 11.4 points apart51.7% to 63.1%
    1. DroidFactory63.1%
    2. Letta CodeLetta59.1%
    3. MuxCoder58.4%
    4. Terminus 2Terminal-Bench57.8%
    5. GooseBlock54.3%
    6. Claude CodeAnthropic52.1%
    7. OpenHandsOpenHands51.9%
    8. OpenCodeAnomaly Innovations51.7%
    Claude Opus 4.5 (20251101) profile
  • GPT-5.1 Codex mini3 agents, 18.5 points apart43.1% to 61.6%
    1. hookeleDmitry Barakhov61.6%
    2. CruxRoam43.1%
    3. Codex CLIOpenAI43.1%
    GPT-5.1 Codex mini profile
  • GPT-5.1 Codex4 agents, 20.9 points apart36.9% to 57.8%
    1. Codex CLIOpenAI57.8%
    2. CruxRoam57.8%
    3. Letta CodeLetta53.5%
    4. Terminus 2Terminal-Bench36.9%
    GPT-5.1 Codex profile
  • GPT-5 (medium)4 agents, 15.7 points apart33.9% to 49.6%
    1. Codex CLIOpenAI49.6%
    2. OpenHandsOpenHands43.8%
    3. Terminus 2Terminal-Bench35.2%
    4. Mini-SWE-AgentPrinceton33.9%
    GPT-5 (medium) profile

Models with results from at least three agents. Terminal-Bench 2.0 (Apache-2.0), compiled by Epoch AI (CC BY 4.0).

Agents

All 40 agents with Terminal-Bench 2.0 results, with who built them and their best result.

  1. NexAU-AHEchina-qijizhifengModels run1Best result84.7%withGPT-5.5Latest resultApr 23, 2026
  2. CapyCapyModels run2Best result83.2%withGPT-5.5Latest resultApr 23, 2026
  3. Codex CLIOpenAIModels run9Best result82.3%withGPT-5.5Latest resultApr 23, 2026
  4. CodexOpenAIModels run1Best result82.0%withGPT-5.5Latest resultApr 23, 2026
  5. ForgeCodeForgeCodeModels run3Best result81.8%withGPT-5.4 (2026-03-05)Latest resultMar 12, 2026
  6. TongAgentsBigaiModels run2Best result80.2%withGemini 3.1 Pro PreviewLatest resultFeb 19, 2026
  7. WOZCODEWOZCODEModels run1Best result80.2%withClaude Opus 4.7Latest resultApr 16, 2026
  8. SageAgentOpenSageModels run2Best result78.4%withGPT-5.3 CodexLatest resultFeb 5, 2026
  9. DroidFactoryModels run5Best result77.3%withGPT-5.3 CodexLatest resultFeb 5, 2026
  10. Meta-HarnessStanford IRISModels run1Best result76.4%withClaude Opus 4.6Latest resultFeb 5, 2026
  11. CodeBrain-1.5Feeling AIModels run2Best result75.8%withGPT-5.3 CodexLatest resultFeb 5, 2026
  12. CodeliakouswModels run1Best result75.7%withGPT-5.3 CodexLatest resultFeb 5, 2026

Sorted by each agent’s best result.

Cost per task

What models cost and how much work they do per task, as the tests measured it.

DeepSWE: score against cost per task

What one task cost on average, against the share of tasks solved. Up and to the left is better value; the line joins the models no cheaper run beats.

  • OpenAI
  • Anthropic
  • Google
  • xAI
  • Z.ai
  • Meta AI
  • Moonshot AI
  • Alibaba
  • Best score at each cost
  • Above-median score, below-median cost

Use the arrow keys to move between models. Press Enter to open the selected model. A table with the same data follows the chart.

DeepSWE: the 12 highest scores with their cost per task
ModelScoreCost per taskOutput tokens per taskAgent steps per task
GPT-6 Astra (xhigh)74.1%$6.5229.6K28.8
Gemini 3.8 Flash (high)73.8%$2.36143K166.3
Claude Opus 5 (max)73.7%$11.84118K99
GPT-6 Astra (high)73.2%$5.7226.5K27.4
GPT-6 Astra (max)73.2%$12.3761.1K28.5
Claude Opus 5 (xhigh)73.2%$9.0791.7K88.7
Claude Opus 5 (high)72.8%$6.0864.2K72.9
GPT-6 Astra (medium)72.8%$4.3820.4K26
GPT-5.6 Sol (max)72.7%$8.3960K61.3
Gemini 3.8 Flash (medium)71.0%$1.97125K147.3
GPT-5.6 Sol (xhigh)70.7%$4.7040.7K44
Claude Fable 5 (xhigh)69.9%$13.4180.4K68.4

Measured and published by Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0). Cost is the API cost of the tokens a run used, at list prices. Tasks and agents differ between tests, so compare costs within one test only.

How long a task models can finish

METR’s time horizons: longer tasks, measured in human time, that models complete.

Task length completed half the time

How long the task takes a skilled person. Longer is better. Log scale.
  1. Claude Mythos Preview17h 25m
  2. Claude Opus 4.611h 59m
  3. Gemini 3.1 Pro Preview6h 24m
  4. GPT-5.2 (high)5h 52m
  5. GPT-5.3 Codex5h 50m
  6. GPT-5.4 (2026-03-05)5h 42m
  7. Claude Opus 4.5 (20251101)4h 53m
  8. Gemini 3 Pro Preview3h 44m
  9. GPT-5.1 Codex Max3h 44m
  10. GPT-5 (high)3h 23m
  11. o3 (2025-04-16)2h
  12. Claude Opus 4.1 (20250805)1h 41m
  13. Claude Opus 4 (20250514)1h 40m
  14. Claude Sonnet 3.71h
  15. o1 (medium)39 min

METR times software tasks with skilled people, then finds the length of task each model completes half the time. The ranges are wide: for Claude Mythos Preview the 95% interval runs from 8h 29m to 55h 4m, and the task length it completes four times out of five is 3h 6m. Only results from the latest method, METR-Horizon-v1.1, are shown. Data: METR time horizons (compiled by Epoch AI, CC BY 4.0).

Models on coding tests

How the models themselves compare when every model runs in the same setup, including releases newer than the agent results.

Arena WebDev

People describe a web page or app; two anonymous models each build it, and people vote for the better working result.

Each model shows its best tested configuration. A test runs every model the same way, so these charts compare models, not agents. Results published by Arena (CC BY 4.0).

All 123 results

Coding score and price

Which models give the most coding ability for their list price.

Arena WebDev against price

Up and to the left is better value: a higher score for a lower list price. The line joins the models no cheaper model beats.

  • Anthropic
  • Alibaba
  • Google
  • OpenAI
  • xAI
  • Z.ai
  • DeepSeek
  • Moonshot AI
  • Xiaomi
  • MiniMax
  • Mistral AI
  • Meta
  • Thinking Machines
  • Tencent
  • Inception
  • StepFun
  • Upstage
  • Best score at each price
  • Above-median score, below-median price

Use the arrow keys to move between models. Press Enter to open the selected model. A table with the same data follows the chart.

Price is the model’s list price per 1M tokens, blending three input tokens with one output token. It is not the cost of a task: an agent can use many tokens on one task. Scores from Arena (CC BY 4.0) and models.dev (MIT).

Frequently asked questions

Answered from the data on this page.

What is an AI coding agent?

A coding agent is the program that turns an AI model into something that can work on code: it gives the model tools to read and edit files, run commands and tests, and keep going until the task is done. Codex CLI, Claude Code, Gemini CLI and OpenHands are examples. The same model can do better or worse depending on the agent around it.

Which coding agent scores highest?

On Terminal-Bench 2.0, NexAU-AHE running GPT-5.5 solves 84.7% of the tasks, the highest result. The best agent from a model maker is Codex CLI with GPT-5.5, at 82.3%.

Does the agent matter as much as the model?

It can. For the 21 models that at least three agents ran, the gap between the best and weakest agent has a median of 13.3 points and reaches 21.8 points for Claude Opus 4.6, from 58.0% with Claude Code to 79.8% with ForgeCode.

Which agent gets the most from GPT-5.5?

NexAU-AHE 84.7%, Capy 83.2%, Codex CLI 82.3%, Codex 82.0%, clnkr 66.1%. NexAU-AHE and Capy are within their standard errors of each other, so treat them as tied.

How much does an AI coding agent cost per task?

It depends on the model, its reasoning effort and the task. DeepSWE, CursorBench, SWE-bench Verified, bash only publish the average API cost per task. On DeepSWE, GPT-6 Astra (xhigh) scores highest at $6.52 a task, while Gemini 3.8 Flash (medium) comes within five points for $1.97. Costs use list prices and differ between tests, so compare them within one test.

How long a task can an AI agent finish?

METR measures the length of software task, in the time a skilled person needs, that a model completes half the time. Claude Mythos Preview leads its measurements at 17h 25m. The uncertainty is large, and the task length a model completes four times out of five is much shorter.

Why are the newest models missing from the agent results?

Agent and model pairs appear here only after the Terminal-Bench 2.0 leaderboard lists them, which can take weeks after a model is released. For newer models, the model tests on this page, such as Arena WebDev, DeepSWE, CursorBench and LiveBench agentic coding, are more current.

Which model is best at agentic coding?

On LiveBench agentic coding, DeepSeek V4.1 Flash (max) scores highest at 77.3%, ahead of Claude Opus 5.5 (max effort) at 71.7%. Other agentic tests on this page rank models differently, so check the test closest to your work.

Which AI is best at building websites and web apps?

On Arena WebDev, where people compare two web apps built from the same request without knowing which model made them, Claude Opus 5.5 (max) scores highest at 1815, ahead of GPT-6 Astra (max) (1788). The scores come from blind votes, so they measure which result people prefer.

Where does this data come from?

Agent results come from the Terminal-Bench 2.0 leaderboard (Apache-2.0, compiled by Epoch AI) and SWE-bench (CC BY-NC 4.0). Model tests come from Arena (CC BY 4.0), LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), SWE-bench, and DeepSWE by Datacurve and CursorBench by Cursor as compiled by Epoch AI; task lengths come from METR, also via Epoch AI. Prices are from models.dev (MIT). Artificials does not run these tests; it matches results to models and calculates the rankings, gaps and charts. Everything refreshes automatically.

Sources and method

Artificials does not run these tests. Rankings, the agent gaps, the matching of results to models and the charts are calculated by Artificials from published results. Some sources allow only non-commercial use; Artificials is a free, non-commercial site.

DeepSWE, CursorBench and METR publish their results without a licence of their own; Epoch AI republishes them under CC BY 4.0, and they are credited to both. Agent names are merged when a source spells one agent several ways, and an entry listed twice keeps its latest run. Whether an agent comes from the model’s maker is worked out from the builder the source names. SWE-bench agent results include only entries that tried each issue once with one model.