Terminal-bench
A benchmark of tasks performed in a real shell — compiling, configuring, debugging, data processing — measuring agentic command-line competence.
Terminal-bench evaluates whether an agent can accomplish tasks a developer would do in a terminal, with the environment as the judge. Anthropic reported Claude Opus 4 at 43.2% at launch in May 2025. It complements SWE-bench, which is about patches to repositories rather than operating a machine.