Benchmarks
How AI models are tested, explained simply, with every benchmark we track and what its score really means.
What it is
A fixed set of questions or tasks, the same for every model, with an automatic way to mark answers.
Why it matters
It lets you compare models on the same test instead of trusting marketing claims.
How to read it
A score is the share of tasks solved. Higher is better, but only within the same test and setup.
The benchmarks
Every test we track, grouped by the skill it measures. Open one for the full story.
Coding
Can it fix real software?SWE-bench Verified
Fix real bugs reported in popular Python projects.
- Tasks
- 500
- Released
- 2024
- Best tracked
- 95.5%
SWE-bench Pro (public set)
Harder, multi-file coding jobs from business and developer software.
- Tasks
- 731
- Released
- 2025
- Best tracked
- 89.9%
Terminal-Bench 2.0
Hard real-world jobs done only by typing commands.
- Tasks
- 89
- Released
- 2025
- Best tracked
- 82.7%
Expert knowledge
Can it answer questions that stump specialists?Puzzle reasoning
Can it figure out a new rule from a few examples?Computer use
Can it operate a real computer like a person?Images and diagrams
Can it understand charts, pictures and diagrams?Customer service agents
Can it follow rules while helping a person?Web research
Can it dig up hard-to-find facts online?Long memory
Can it find the right detail in a huge conversation?Understand the numbers
A score is a sample
A benchmark is a fixed set of tasks with an automatic grader. The score is the share of tasks the model gets right. Run it and watch: the same model can land on different scores.
In plain wordsLike a driving test with 20 manoeuvres: a good driver can still fail one on a bad day. More manoeuvres give a fairer picture.
One run is one sample. Small sets wobble more.
pass@1 and pass@k
Coding benchmarks often let a model try several times. pass@k is the chance that at least one of k tries is correct. Labs estimate it from n samples, c of which were right.
In plain wordsLike letting a student retake a quiz five times and keeping the best try. Useful, but not the same as getting it right first time.
pass@5 always looks better than pass@1. Only compare like with like.
How sure is a score?
With N questions and a score p, the standard error tells you how much the score would move if you drew a fresh set of questions. The 95% interval is about two standard errors either side.
In plain wordsLike polling 200 people before an election: the result is close to the truth, but a small lead may be luck.
On a 200-question test, a two-point lead is often inside the noise.
Same name, different test
Two numbers with the same benchmark name can come from different setups. Before comparing, check these five things.
In plain wordsLike comparing two runners when one wore spikes and the other ran uphill. Same race name, different conditions.
Compare within one column and one setup, never across.
Saturation and contamination
When top models all score near 100%, a benchmark stops telling them apart (saturation). If test questions leak into training data, scores rise without real skill (contamination). Newer, harder or private sets are the answer.
In plain wordsLike a school exam everyone has already seen the answers to: high marks stop meaning much.
A high score on an old public test means less than it used to.
Points per dollar
The best model for you is rarely the top scorer. Divide the score by what one request costs to see value, then check that the cheaper model is good enough for your task.
In plain wordsLike choosing a car: the fastest one is not the best buy if a cheaper one gets you there just as well.
Pick the cheapest model that clears your bar, not the highest bar.
Ready to compare real models?
Open Compare