DictionaryEvaluation See it in actionBenchmarks explained
SWE-bench
A test where AI fixes real bugs from open-source projects, checked by running each project's own tests.
Definition
A coding benchmark of 2,294 real GitHub issues from 12 popular Python repositories (Jimenez et al., ICLR 2024): the model edits the codebase to resolve each issue, and the fix is checked with the repository's tests. SWE-bench Verified is a 500-task subset whose descriptions, tests, and solvability were checked by human annotators, created by the SWE-bench team with OpenAI. Other variants include Lite, Multimodal, Multilingual, and the harder SWE-bench Pro. At launch, Claude 2 resolved 1.96% of issues.
Example
Given a Django issue about a wrong date format, the model edits the repo; it scores only if the hidden tests for that fix pass.