From one question to a ranking
How AI benchmarks work
How a test turns answers into a score, how sure that score is, what happens when a test leaks, and how many tests become one ranking. With real results from the tests on this site.
Scroll to beginChapter 1 of 7
One question
Every benchmark starts small: a question with a known answer.
A benchmark is a set of questions with known answers. This one is multiple choice: four options, one right.
Weighed
The model’s choice comes from probabilities, like every token it writes. Here they fall on four letters.
Answered
It answers B. Many tests let a model reason first, then read only its final answer.
Graded
The answer is checked against the key: one point if it matches, none if it does not. No judgement is needed, which makes multiple choice cheap to grade.
Chapter 2 of 7
Hundreds of questions
One question proves little. A test asks hundreds.
GPQA Diamond has 198 graduate-level questions in biology, chemistry and physics, written by experts to be hard to look up.
GPQA: Rein et al., 2023
Answered
An illustrative model answers every question and gets 159 right, one question at a time.
Scored
159 of 198 is 80.3%. With four choices each, guessing alone scores about 25%: the floor is not zero.
Chapter 3 of 7
How sure is a score?
A score is a measurement, and measurements have error bars.
Think of the 198 questions as a sample of every question the test could have asked. Another sample of the same kind would give a slightly different score.
Error bars
With 198 questions and a score of 80%, the standard error is about 2.8 points: 95% of such tests would land between 74.8% and 85.8%.
Error bars for evaluations: Miller, 2024
Part two: other tests
Not every test is a quiz.
Chapter 4 of 7
Four ways to grade
What counts as right depends on the test.
Multiple choice is graded by matching one letter to the key.
An exact answer
Maths tests often want a single number and accept only that number. In the AIME exam, every answer is a whole number from 0 to 999, so guessing almost never works.
Working code
SWE-bench gives a model a real issue from an open-source Python project. The model changes the code, and scores only if the project’s own tests pass.
SWE-bench: Jimenez et al., 2023
People’s votes
Some qualities have no answer key. Arena shows a person two anonymous answers to their own question, and they pick the better one.
Chatbot Arena: Chiang et al., 2024
Chapter 5 of 7
Votes become ratings
Millions of one-to-one comparisons, turned into one ranking.
Each vote says only that one model beat another, once.
Ratings
A statistical model, Bradley–Terry, finds the scores that best explain every vote. The further apart two scores, the more often the higher one is expected to win.
Chapter 6 of 7
When the test leaks
A model that has seen the answers is not showing skill.
A test’s questions are meant to stay out of a model’s training data.
Leaked
But a published test lives on the web, and training data is gathered from the web. A copy can slip in: this is called contamination.
Inflated
A model that has seen a question can answer it from memory. Its score on those questions can rise without any more skill.
Keeping tests fresh
LiveBench adds new questions every month from recent sources. Other tests keep a set private, and researchers search training data for copies.
LiveBench: White et al., 2024
Part three: many tests
So which model is best?
Chapter 7 of 7
Many tests, one scale
No single test can rank every model. Together, they can.
An easy test tells weaker models apart but leaves the best tied near 100%. A hard test does the opposite.
One scale
Fit the tests together: each gets a difficulty, each model a capability, and a curve links the two. Tests at every level then place all the models on one scale.
The capability index
The Epoch Capabilities Index combines more than fifty benchmarks this way. Its scale is set so that Claude 3.5 Sonnet scores 130 and GPT-5 150: compare its numbers, don’t read them alone. Today Claude Opus 5.5 leads, and 8 of the top eight reach into its range.
Before you trust a ranking
Read the fine print.
- Which test is it, and what does it measure?
- How many questions does it have?
- Are the differences bigger than the error bars?
- Could its questions have leaked into training data?
- Who ran it, when, and with which settings?
Sources
The example question, the illustrative model and the leaked scores are written for this page and marked where they appear. Real results come from Epoch AI (CC BY 4.0) and Arena (CC BY 4.0). More on reading results: A benchmark is a signal and Benchmark contamination.
- Rein et al., 2023: GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Miller, 2024: Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Jimenez et al., 2023: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Chiang et al., 2024: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- White et al., 2024: LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- Epoch AI: Epoch Capabilities Index