Time horizon (50%-task-completion)
The length of task, in human-expert time, an agent completes with 50% success.
METR's time-horizon metric fits a logistic curve of agent success against how long each task takes a skilled human, and reports the task length at which success is 50% (also 80%). It converts scattered benchmark scores into a single, comparable capability number. The 50% horizon has doubled roughly every 7 months since 2019; GPT-5 measured about 2h17m in August 2025. Above ~16 hours the current task suite cannot measure reliably.