All makers

AI model maker

Anthropic

40 models on Artificials, 22 with published test results.

RSS feed
Best on the capability index167.3Claude Opus 5.5, #1 of 274
Top score on10 testsof the 27 its models are tested on
Newest releaseClaude Sonnet 5.5Sep 28, 2026
Models4022 tested, 0 with open weights

Where Anthropic stands

Its best model’s rank among all models tested on each skill area’s main test, the same tests model pages use. A longer bar means a better rank.

  • OverallClaude Opus 5.5 on Epoch Capabilities Index
    #1 of 274
  • CodingClaude Sonnet 5.5 (max effort) on LiveBench coding
    #1 of 63
  • MathClaude Fable 5 (high) on Mock AIME 2024–2025
    #1 of 198
  • KnowledgeClaude Opus 5.5 (max) on SimpleQA Verified
    Top 6%
  • AgentsClaude Opus 5.5 (max effort) on LiveBench agentic coding
    Top 4%
  • LanguageClaude Fable 5 (max effort) on LiveBench language
    #1 of 63
  • Data analysisClaude Fable 5 (max effort) on LiveBench data analysis
    Top 8%
  • ReasoningClaude Sonnet 5.5 (max) on GPQA Diamond
    Top 1%

On the capability index

Its best 12 of 26 models on the Epoch Capabilities Index, each with its rank among all 274 models. Claude Opus 5.5 holds the top score, 167.3.

Every test

Its best result on each published test, and that result’s rank among every model tested. Scores from different tests are never combined.

TestBest modelScoreRank
Epoch Capabilities IndexClaude Opus 5.5167.3#1 of 274
Mock AIME 2024–2025Claude Fable 5 (high)100.0%#1 of 198
Arena hard promptsClaude Opus 5.5 (high)1534#2 of 389
Arena instruction followingClaude Opus 5.5 (high)1517#2 of 389
Arena textClaude Opus 4.6 (high)1505#2 of 389
Arena creative writingClaude Opus 5.5 (high)1516#2 of 387
Arena codingClaude Fable 5 (high)1552#2 of 384
Arena mathClaude Fable 5 (high)1522#2 of 372
Arena WebDevClaude Opus 5.5 (max)1815#1 of 123
GPQA DiamondClaude Sonnet 5.5 (max)95.6%#2 of 222
LiveBench codingClaude Sonnet 5.5 (max effort)91.4%#1 of 63
LiveBench languageClaude Fable 5 (max effort)90.7%#1 of 63
LiveBench mathematicsClaude Opus 5.5 (max effort)97.1%#1 of 63
LiveBench overallClaude Fable 5.1 (max effort)83.4%#1 of 63
SWE-bench Verified, bash onlyClaude Opus 4.5 (high)76.8%#1 of 42
SWE-bench VerifiedClaude Opus 4.7 (max)83.5%#1 of 32
LiveBench agentic codingClaude Opus 5.5 (max effort)71.7%#2 of 63
FrontierMath Tiers 1–3Claude Opus 5.5 (max)91.2%#3 of 81
FrontierMath Tier 4Claude Opus 5.5 (max)95.0%#3 of 63
LiveBench reasoningClaude Opus 5.5 (max effort)92.2%#3 of 63
MATH Level 5Claude Sonnet 4.5 (20250929) (32K)97.7%#5 of 97
SimpleQA VerifiedClaude Opus 5.5 (max)72.2%#4 of 77
CursorBenchClaude Opus 5.5 (max)57.8%#1 of 14
LiveBench data analysisClaude Fable 5 (max effort)80.5%#5 of 63
Chess puzzlesClaude Fable 5.1 (max)47.0%#13 of 141
LiveBench instructionsClaude Fable 5 (max effort)75.8%#6 of 63
DeepSWEClaude Opus 5 (max)73.7%#3 of 26

Every model

Newest first. The price is per 1M tokens, blended three input tokens to one output token: Anthropic’s own listing, or else the middle price across hosts.

ModelReleasedContextPriceCapability index
Claude Sonnet 5.5Sep 28, 20261M tokens$4.00165.0
Claude Opus 5.5Sep 22, 20261M tokens$8.00167.3
Claude Opus 5.5 FastSep 22, 20261M tokens$16.00Not on the index
Claude Fable 5.1Sep 1, 20261M tokens$20.00164.7
Claude Opus 5Jul 24, 20261M tokens$10.00162.8
Claude Opus 5 (fast)Jul 23, 20261M tokens$20.00Not on the index
Claude Sonnet 5Jun 29, 20261M tokens$4.00156.2
Claude Fable LatestJun 9, 20261M tokens$20.00Not on the index
Claude Mythos 5Jun 9, 20261M tokens$20.00Not on the index
Claude Fable 5Jun 7, 20261M tokens$20.00162.1
Claude Opus 4.8May 28, 20261M tokens$10.00158.2
Claude Opus 4.8 FastMay 28, 20261M tokens$20.00Not on the index
Claude Opus 4.7Apr 14, 20261M tokens$10.00156.3
Claude Haiku LatestMar 29, 2026200K tokens$2.00Not on the index
Claude Opus LatestMar 29, 20261M tokens$8.00Not on the index

Models, prices, context windows and release dates from models.dev (MIT). Test results from LiveBench (Apache-2.0), Epoch AI (CC BY 4.0), Datacurve DeepSWE, compiled by Epoch AI (CC BY 4.0), Cursor CursorBench, compiled by Epoch AI (CC BY 4.0), Arena (CC BY 4.0) and SWE-bench (CC BY-NC 4.0). A model counts here when Anthropic sells it, several hosts list it or a test covers it. Each figure is the source’s own; nothing is combined into a new score.