Blog

AI Agent Glossary · AI Coding Tools Glossary

Benchmark saturation

The point at which top models score near the ceiling of a benchmark so it can no longer distinguish them; HumanEval and SWE-bench Verified are saturated.

SWE-bench gained 67 points in a single year (2023→2024, per Stanford's AI Index) and Verified went from ~60% to near 100% during 2025. Saturation drives the churn of benchmarks — HumanEval → SWE-bench → SWE-bench Verified → SWE-bench Pro / Terminal-bench — and is why a benchmark's age matters as much as the score.

Source: AI Index Report 2025 — Stanford HAI

See the numbers behind this term: AI Agent Statistics 2026 · AI Coding Tools Statistics 2026.