@tokalator4.2.0…installs…downloads

Benchmarks

How AI models are tested, explained simply, with every benchmark we track and what its score really means.

What it is

A fixed set of questions or tasks, the same for every model, with an automatic way to mark answers.

Why it matters

It lets you compare models on the same test instead of trusting marketing claims.

How to read it

A score is the share of tasks solved. Higher is better, but only within the same test and setup.

The benchmarks

Every test we track, grouped by the skill it measures. Open one for the full story.

Images and diagrams

Can it understand charts, pictures and diagrams?

Understand the numbers

01

A score is a sample

A benchmark is a fixed set of tasks with an automatic grader. The score is the share of tasks the model gets right. Run it and watch: the same model can land on different scores.

In plain wordsLike a driving test with 20 manoeuvres: a good driver can still fail one on a bad day. More manoeuvres give a fairer picture.

One run is one sample. Small sets wobble more.

This run—
True skill70%
Last runs
02

pass@1 and pass@k

Coding benchmarks often let a model try several times. pass@k is the chance that at least one of k tries is correct. Labs estimate it from n samples, c of which were right.

In plain wordsLike letting a student retake a quiz five times and keeping the best try. Useful, but not the same as getting it right first time.

pass@5 always looks better than pass@1. Only compare like with like.

pass@130.0%
pass@587.1%
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
03

How sure is a score?

With N questions and a score p, the standard error tells you how much the score would move if you drew a fresh set of questions. The 95% interval is about two standard errors either side.

In plain wordsLike polling 200 people before an election: the result is close to the truth, but a small lead may be luck.

On a 200-question test, a two-point lead is often inside the noise.

A
72.3% to 83.7%
B
74.5% to 85.5%

A 2-point gap on 200 questions could be noise.

04

Same name, different test

Two numbers with the same benchmark name can come from different setups. Before comparing, check these five things.

In plain wordsLike comparing two runners when one wore spikes and the other ran uphill. Same race name, different conditions.

Compare within one column and one setup, never across.

0 of 5 checked. Tap each one you have confirmed.

05

Saturation and contamination

When top models all score near 100%, a benchmark stops telling them apart (saturation). If test questions leak into training data, scores rise without real skill (contamination). Newer, harder or private sets are the answer.

In plain wordsLike a school exam everyone has already seen the answers to: high marks stop meaning much.

A high score on an old public test means less than it used to.

Illustrative curve, not real data
100%0%saturatedtime since release →
06

Points per dollar

The best model for you is rarely the top scorer. Divide the score by what one request costs to see value, then check that the cheaper model is good enough for your task.

In plain wordsLike choosing a car: the fastest one is not the best buy if a cheaper one gets you there just as well.

Pick the cheapest model that clears your bar, not the highest bar.

Model A
Model B
Model A
137 points per $
Model B
617 points per $

Model B gives 4.5× more score per dollar.

Ready to compare real models?

Open Compare