@tokalator4.2.0…installs…downloads
All benchmarks
Coding

Terminal-Bench 2.0

Hard real-world jobs done only by typing commands.

Tasks
89
Released
Nov 2025
Made by
Stanford University and Laude Institute, with nearly 100 contributors
Score
Percent of tasks completed
In plain words

Like giving a technician remote access to a computer with only a command line, then checking whether the job got done.

Example taskSource

Rewrite an old COBOL program in Python so it produces exactly the same output as the original.

A written instruction inside a sealed computer environment. The AI types commands until it thinks the job is done.

How it is marked

Automated tests check the final state of the computer. Each task also has a human-written solution proving it is doable.

Scores we track

Each number is the provider's own published result. Tap one for its source.

  1. 1GPT-5.582.7%
  2. 2GPT-5.3-Codex77.3%
  3. 3GPT-5.475.1%
  4. 4Claude Opus 4.769.4%
  5. 5Gemini 3.1 Pro68.5%
  6. 6Claude Opus 4.665.4%
  7. 7GPT-5.4 mini60%
  8. 8Claude Opus 4.559.3%
  9. 9Claude Sonnet 4.659.1%
  10. 10Gemini 3 Flash47.6%
  11. 11GPT-5.4 nano46.3%
  12. 12Gemini 2.5 Pro32.6%

Why it matters

Coding assistants like Claude Code and Codex CLI work through the terminal, so this mirrors how they are actually used.

Good to know

  • Scores depend on the agent software wrapped around the model, not only the model.
  • Newer versions now exist (2.1, 3.0 and 4.0), so check which one a score refers to.
  • Tasks span software, system administration, data science and more, not just coding.