@tokalator4.2.0…installs…downloads
Dictionary
Evaluation

Benchmark

A standard exam for AI models: the same questions for every model, graded automatically, so you can compare scores fairly.

Definition

A fixed set of tasks with an automatic grader, run the same way on every model so results can be compared. The score is the share of tasks passed. Scores are only comparable within the same variant and setup (attempts, tools, effort, harness), and independent evaluators such as Epoch AI track results across models over time.

Example

Model A solves 410 of 500 SWE-bench Verified tasks, scoring 82%. Comparing that with another lab's pass@5 score would be misleading.

Appears in

See it in actionBenchmarks explained