Blog

AI Coding Tools Glossary

SWE-bench Verified

A 500-task subset of SWE-bench screened by human annotators to remove ambiguous or under-specified issues.

OpenAI and the SWE-bench authors had engineers review tasks and keep only those with clear problem statements and fair tests, producing a cleaner 500-problem set in August 2024. Scores rose from ~60% at the start of 2025 to near saturation by the end of the year, which is why Verified is no longer discriminative between frontier models. The benchmark is public and Python-only, so contamination and overfitting are concerns.

Source: AI Index Report 2026 — Stanford HAI

See the numbers behind this term: AI Coding Tools Statistics 2026.