SWE-bench Verified
A 500-task subset of SWE-bench screened by human annotators to remove ambiguous or under-specified issues.
OpenAI and the SWE-bench authors had engineers review tasks and keep only those with clear problem statements and fair tests, producing a cleaner 500-problem set in August 2024. Scores rose from ~60% at the start of 2025 to near saturation by the end of the year, which is why Verified is no longer discriminative between frontier models. The benchmark is public and Python-only, so contamination and overfitting are concerns.