@tokalator4.2.0…installs…downloads
All benchmarks
Coding

SWE-bench Verified

Fix real bugs reported in popular Python projects.

Tasks
500
Released
Aug 2024
Made by
SWE-bench team (Princeton University) with OpenAI
Score
Percent of tasks resolved
In plain words

Like handing a new hire a real bug report and the codebase, then checking whether their fix passes the project's own tests.

Example taskSource

In the original SWE-bench paper, a scikit-learn issue reports a data leak in gradient boosting when warm start is used; the model must patch it.

A real GitHub issue plus a snapshot of the code. The model must write a code change that resolves it.

How it is marked

The project's tests run on the changed code. Tests tied to the issue must now pass, and previously passing tests must still pass.

Scores we track

Each number is the provider's own published result. Tap one for its source.

  1. 1Claude Mythos 595.5%
  2. 2Claude Fable 595%
  3. 3Claude Opus 4.888.6%
  4. 4Claude Opus 4.787.6%
  5. 5Claude Sonnet 585.2%
  6. 6Claude Opus 4.580.9%
  7. 7Claude Opus 4.680.8%
  8. 8Gemini 3.1 Pro80.6%
  9. 9Claude Sonnet 4.679.6%
  10. 10Gemini 3 Flash78%
  11. 11Claude Haiku 4.573.3%
  12. 12Gemini 2.5 Pro59.6%

Why it matters

It is the most widely quoted test of whether AI can do everyday bug-fixing work in real projects.

Good to know

  • Python only, drawn from 12 open-source projects.
  • Humans checked each task was clear and solvable, which removed many flawed ones.
  • The SWE-bench Pro authors note many Verified tasks are fairly small fixes.