@tokalator4.2.0…installs…downloads
Dictionary
Evaluation

Terminal-Bench

A test of whether an AI agent can get real jobs done in a computer terminal, like a sysadmin practical exam.

Definition

A benchmark of hard, realistic tasks that agents must complete in a command-line environment, hosted by Stanford and the Laude Institute with Harbor. Terminal-Bench 2.0 has 89 tasks, each with its own environment, a human-written solution, and tests that verify the result; frontier models and agents scored below 65% when its paper was published in 2026. Later versions continue the series.

Example

A task might ask the agent to compile an old C project with missing dependencies; tests then check that the binary builds and runs.

Appears in

See it in actionBenchmarks explained