All benchmarksCoding In plain words Example taskSource
Terminal-Bench 2.0
Hard real-world jobs done only by typing commands.
- Tasks
- 89
- Released
- Nov 2025
- Made by
- Stanford University and Laude Institute, with nearly 100 contributors
- Score
- Percent of tasks completed
Like giving a technician remote access to a computer with only a command line, then checking whether the job got done.
Rewrite an old COBOL program in Python so it produces exactly the same output as the original.
A written instruction inside a sealed computer environment. The AI types commands until it thinks the job is done.How it is marked
Automated tests check the final state of the computer. Each task also has a human-written solution proving it is doable.
Scores we track
Each number is the provider's own published result. Tap one for its source.
Why it matters
Coding assistants like Claude Code and Codex CLI work through the terminal, so this mirrors how they are actually used.
Good to know
- Scores depend on the agent software wrapped around the model, not only the model.
- Newer versions now exist (2.1, 3.0 and 4.0), so check which one a score refers to.
- Tasks span software, system administration, data science and more, not just coding.