From one question to a ranking

How AI benchmarks work

How a test turns answers into a score, how sure that score is, what happens when a test leaks, and how many tests become one ranking. With real results from the tests on this site.

Scroll to begin

Chapter 1 of 7

One question

Every benchmark starts small: a question with a known answer.

A benchmark is a set of questions with known answers. This one is multiple choice: four options, one right.

Weighed

The model’s choice comes from probabilities, like every token it writes. Here they fall on four letters.

Answered

It answers B. Many tests let a model reason first, then read only its final answer.

Graded

The answer is checked against the key: one point if it matches, none if it does not. No judgement is needed, which makes multiple choice cheap to grade.

Chapter 2 of 7

Hundreds of questions

One question proves little. A test asks hundreds.

GPQA Diamond has 198 graduate-level questions in biology, chemistry and physics, written by experts to be hard to look up.

GPQA: Rein et al., 2023

Answered

An illustrative model answers every question and gets 159 right, one question at a time.

Scored

159 of 198 is 80.3%. With four choices each, guessing alone scores about 25%: the floor is not zero.

Near the ceiling

The best result on GPQA Diamond today is 95.8%, by GPT-6 Astra (max). When the best models crowd near 100%, a test can no longer tell them apart.

Results: Epoch AI (CC BY 4.0)

Chapter 3 of 7

How sure is a score?

A score is a measurement, and measurements have error bars.

Think of the 198 questions as a sample of every question the test could have asked. Another sample of the same kind would give a slightly different score.

Error bars

With 198 questions and a score of 80%, the standard error is about 2.8 points: 95% of such tests would land between 74.8% and 85.8%.

Error bars for evaluations: Miller, 2024

Real results

Epoch AI publishes a standard error with each result. On GPQA Diamond, 8 of the top eight reach into the leader’s range: too close to rank with confidence.

Results: Epoch AI (CC BY 4.0)

Part two: other tests

Not every test is a quiz.

Chapter 4 of 7

Four ways to grade

What counts as right depends on the test.

Multiple choice is graded by matching one letter to the key.

An exact answer

Maths tests often want a single number and accept only that number. In the AIME exam, every answer is a whole number from 0 to 999, so guessing almost never works.

Working code

SWE-bench gives a model a real issue from an open-source Python project. The model changes the code, and scores only if the project’s own tests pass.

SWE-bench: Jimenez et al., 2023

People’s votes

Some qualities have no answer key. Arena shows a person two anonymous answers to their own question, and they pick the better one.

Chatbot Arena: Chiang et al., 2024

Chapter 5 of 7

Votes become ratings

Millions of one-to-one comparisons, turned into one ranking.

Each vote says only that one model beat another, once.

Ratings

A statistical model, Bradley–Terry, finds the scores that best explain every vote. The further apart two scores, the more often the higher one is expected to win.

Real results

Arena publishes a 95% range with each score. Among its text leaders, 1 of the top eight reach into the leader’s.

Results: Arena (CC BY 4.0)

Chapter 6 of 7

When the test leaks

A model that has seen the answers is not showing skill.

A test’s questions are meant to stay out of a model’s training data.

Leaked

But a published test lives on the web, and training data is gathered from the web. A copy can slip in: this is called contamination.

Inflated

A model that has seen a question can answer it from memory. Its score on those questions can rise without any more skill.

Keeping tests fresh

LiveBench adds new questions every month from recent sources. Other tests keep a set private, and researchers search training data for copies.

LiveBench: White et al., 2024

Part three: many tests

So which model is best?

Chapter 7 of 7

Many tests, one scale

No single test can rank every model. Together, they can.

An easy test tells weaker models apart but leaves the best tied near 100%. A hard test does the opposite.

One scale

Fit the tests together: each gets a difficulty, each model a capability, and a curve links the two. Tests at every level then place all the models on one scale.

The capability index

The Epoch Capabilities Index combines more than fifty benchmarks this way. Its scale is set so that Claude 3.5 Sonnet scores 130 and GPT-5 150: compare its numbers, don’t read them alone. Today Claude Opus 5.5 leads, and 8 of the top eight reach into its range.

Epoch AI (CC BY 4.0)

Before you trust a ranking

Read the fine print.

  • Which test is it, and what does it measure?
  • How many questions does it have?
  • Are the differences bigger than the error bars?
  • Could its questions have leaked into training data?
  • Who ran it, when, and with which settings?

Sources

The example question, the illustrative model and the leaked scores are written for this page and marked where they appear. Real results come from Epoch AI (CC BY 4.0) and Arena (CC BY 4.0). More on reading results: A benchmark is a signal and Benchmark contamination.

  1. Rein et al., 2023: GPQA: A Graduate-Level Google-Proof Q&A Benchmark
  2. Miller, 2024: Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
  3. Jimenez et al., 2023: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
  4. Chiang et al., 2024: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
  5. White et al., 2024: LiveBench: A Challenging, Contamination-Limited LLM Benchmark
  6. Epoch AI: Epoch Capabilities Index