DictionaryEvaluation See it in actionBenchmarks explained
Terminal-Bench
A test of whether an AI agent can get real jobs done in a computer terminal, like a sysadmin practical exam.
Definition
A benchmark of hard, realistic tasks that agents must complete in a command-line environment, hosted by Stanford and the Laude Institute with Harbor. Terminal-Bench 2.0 has 89 tasks, each with its own environment, a human-written solution, and tests that verify the result; frontier models and agents scored below 65% when its paper was published in 2026. Later versions continue the series.
Example
A task might ask the agent to compile an old C project with missing dependencies; tests then check that the binary builds and runs.