Every record, year by year

The AI race

274 models on one scale since Feb 2023. Follow the records, the makers who set them, open models closing in, and what a score costs.

Scroll to begin

Chapter 1 of 6

Every dot is a model

One scale for every model: Epoch AI’s Capabilities Index.

The index combines results from more than fifty benchmarks into one scale. Each model on it becomes a dot: when it came out, across; how it scored, up.

Capability index: Epoch AI (CC BY 4.0)

The field fills in

274 models, from Feb 2023 to Sept 2026. The scale has no top: compare the dots, and don’t read them as percentages.

Who made them

Six makers have the strongest models: Anthropic, OpenAI, Moonshot AI, Google DeepMind, Meta AI and xAI. Every other maker’s models are grey.

The leaders today

  1. Claude Opus 5.5, Anthropic: 167.3
  2. GPT-6 Astra, OpenAI: 166.4
  3. GPT-6.1 Sol, OpenAI: 166.1
  4. Claude Sonnet 5.5, Anthropic: 165.0
  5. Claude Fable 5.1, Anthropic: 164.7

Chapter 2 of 6

The record changes hands

Each time a model beat every earlier one, the record moved up.

The first record on the index is Llama 65B, from Meta AI, in Feb 2023, at 110.2.

Up and up

It has been broken 22 times since, and passed from one maker to another 12 times.

Who held it

23 records by 4 makers: OpenAI 15, Anthropic 5, Google 2 and Meta AI 1. The latest is Claude Opus 5.5, at 167.3.

Capability index: Epoch AI (CC BY 4.0)

Chapter 3 of 6

Each maker’s climb

The six strongest makers, each against its own best.

Each line follows one maker’s best model to date.

Everyone climbs

Every line only steps up: each step is a model better than all that maker’s earlier ones.

A close finish

Today the six best models are within 10.9 points: from Anthropic at 167.3 to xAI at 156.4.

Chapter 4 of 6

Open weights close in

Models anyone can download, against those only their makers run.

Open-weights models can be downloaded and run by anyone. The best of them today scores 157.4; the best closed model, 167.3.

How far behind

Kimi K3 reached 157.4 in Jul 2026. A closed model first scored that much 4 months earlier: GPT-5.4 Pro.

Capability index: Epoch AI (CC BY 4.0)

Part two: price

Getting better is half the story.

Chapter 5 of 6

The price of a score

What the cheapest model at a given level costs, and how fast that falls.

Each model is placed at its release date with today’s list price per million tokens, blended three input tokens to one output token: not its launch price. The line follows the cheapest model at or above one level of the index; the scale falls by ten times at each gridline.

Prices: models.dev (MIT)

477 times cheaper

Of the models released by Dec 2024, the cheapest scoring 140 or more is o1, listed at $26.25. Of those released by Jul 2026, it is Qwen3.7 Flash, at $0.055.

At every level

Each line is one level of the index, from 120 up. Every one only steps down: each point is a new cheapest model at that level.

Chapter 6 of 6

Where the race stands

The highest scores today, and how sure the index is of each.

  1. Claude Opus 5.5, Anthropic: 167.3
  2. GPT-6 Astra, OpenAI: 166.4
  3. GPT-6.1 Sol, OpenAI: 166.1
  4. Claude Sonnet 5.5, Anthropic: 165.0
  5. Claude Fable 5.1, Anthropic: 164.7
  6. Claude Opus 5, Anthropic: 162.8
  7. GPT-6 Sol, OpenAI: 162.7
  8. GPT-5.5 Pro, OpenAI: 162.1

Too close to call

Each score comes with a confidence interval. 8 models reach into Claude Opus 5.5’s: on this index, their order could change as more results arrive.

Capability index: Epoch AI (CC BY 4.0)

The race goes on

Who leads next?

Claude Opus 5.5, from Anthropic, holds the record at 167.3.

This page is drawn from the same data as the rest of the site, so new models appear here as they join the index.

Sources and method

Scores are the Epoch Capabilities Index, with release dates and weight access as Epoch AI records them (CC BY 4.0). Prices are today’s list prices per million tokens from the models.dev catalog (MIT), placed at each model’s release date, not launch prices: the maker’s own listing, or the middle listing across hosts, blended three input tokens to one output token.

A record is a model that scored higher than every model released before it. Artificials draws these charts from the published data; it does not test models itself.