Public–private benchmark gap
The drop in resolve rate between open-source benchmark tasks and tasks from unpublished commercial repositories.
SWE-bench Pro reports both. GPT-5 fell from 23.3% (public) to 14.9% (private), Claude Opus 4.1 from 22.7% to 17.8% — a 22–36% relative drop. Task and repository differences contribute alongside contamination, but the direction is consistent: expect lower scores on your own codebase than on the leaderboard.