SWE-bench Pro
Scale AI's harder successor to SWE-bench: 1,865 long-horizon tasks across 41 repositories, with public, held-out and private commercial subsets.
SWE-bench Pro tasks require on average 107 changed lines across 4.1 files and include languages beyond Python; a private set of unpublished commercial repositories measures performance on code models have never seen. At release in September 2025 the best model scored ~23%; by September 2026 the top public score is 61.5%. Every model scores lower on the private set (GPT-5: 23.3% → 14.9%), the clearest public evidence of a benchmark-to-real-work gap.