Test-time compute (parallel sampling)
Running a model multiple times on the same task and selecting the best attempt, raising benchmark scores at extra cost.
Vendors report two numbers for the same model: single attempt and 'with parallel test-time compute'. Anthropic reports Claude Opus 4 at 72.5% on SWE-bench Verified single-run and 79.4% with parallel sampling and selection. When comparing benchmark claims, check which regime is reported; leaderboards like SWE-bench Pro require a single trajectory.