All benchmarksCoding In plain words Example taskSource
SWE-bench Verified
Fix real bugs reported in popular Python projects.
- Tasks
- 500
- Released
- Aug 2024
- Made by
- SWE-bench team (Princeton University) with OpenAI
- Score
- Percent of tasks resolved
Like handing a new hire a real bug report and the codebase, then checking whether their fix passes the project's own tests.
In the original SWE-bench paper, a scikit-learn issue reports a data leak in gradient boosting when warm start is used; the model must patch it.
A real GitHub issue plus a snapshot of the code. The model must write a code change that resolves it.How it is marked
The project's tests run on the changed code. Tests tied to the issue must now pass, and previously passing tests must still pass.
Scores we track
Each number is the provider's own published result. Tap one for its source.
Why it matters
It is the most widely quoted test of whether AI can do everyday bug-fixing work in real projects.
Good to know
- Python only, drawn from 12 open-source projects.
- Humans checked each task was clear and solvable, which removed many flawed ones.
- The SWE-bench Pro authors note many Verified tasks are fairly small fixes.