All benchmarksCoding In plain words Example taskSource
SWE-bench Pro (public set)
Harder, multi-file coding jobs from business and developer software.
- Tasks
- 731
- Released
- Sep 2025
- Made by
- Scale AI
- Score
- Resolve rate: percent of tasks solved
A tougher sequel to SWE-bench: instead of quick fixes, these are jobs that could take a professional engineer hours or days.
Add Google Books as a backup source of book details for Open Library's import tool when other sources have gaps.
A rewritten issue with clear requirements in a real repository. The model must change the code to deliver it.How it is marked
Human-reviewed tests check the issue is solved and that nothing that used to work is now broken.
Scores we track
Each number is the provider's own published result. Tap one for its source.
- 1Claude Opus 5.589.9%
- 2Claude Sonnet 5.581.3%
- 3Claude Fable 5.181.2%
- 4Claude Mythos 5.181.2%
- 5Claude Mythos 580.3%
- 6Claude Fable 580%
- 7Claude Opus 579.2%
- 8Claude Opus 4.869.2%
- 9Claude Haiku 5.564.8%
- 10Claude Opus 4.764.3%
- 11Claude Sonnet 563.2%
- 12Gemini 3.6 Flash58.7%
- 13GPT-5.558.6%
- 14GPT-5.457.7%
- 15GPT-5.3-Codex56.8%
- 16GPT-5.4 mini54.4%
- 17Gemini 3.1 Pro54.2%
- 18Gemini 3.5 Flash53.9%
- 19Claude Opus 4.653.4%
- 20GPT-5.4 nano52.4%
- 21Gemini 3 Flash48.4%
Why it matters
It tests longer, messier work closer to real engineering jobs, where simpler coding tests are near their ceiling.
Good to know
- The full benchmark has 1,865 tasks; only the 731-task public set is open to everyone.
- Reference fixes average about 107 lines across 4 files, so tasks are much larger than SWE-bench Verified.
- Covers Python, Go, JavaScript and TypeScript only.
- Public repositories use copyleft licenses, chosen to reduce the chance models saw the answers in training.