@tokalator4.2.0…installs…downloads
Dictionary
Evaluation

pass@k

The chance an AI solves a task in at least one of k tries, like keeping a student's best attempt.

Definition

The probability that at least one of k sampled attempts at a task is correct, introduced with the HumanEval benchmark by Chen et al. (2021). It is estimated from n samples with c correct as 1 - C(n-c, k) / C(n, k). pass@1 measures first-try accuracy; a higher k always looks better, so compare scores only at the same k. Codex solved 28.8% of HumanEval problems in one try and 70.2% with 100 samples per problem.

Example

A model passes 3 of 10 samples on a coding task: pass@1 is 0.3, while pass@5 is about 0.92.

See it in actionBenchmarks explained