OSWorld
The reference benchmark for computer-use agents: 369 real tasks on real operating systems.
OSWorld runs agents inside actual Ubuntu, Windows and macOS VMs and scores them with 134 execution-based evaluators (checking files, application state) rather than text matching. Humans complete 72.36% of tasks. OSWorld-Verified (July 2025) fixed community-reported task issues and cut evaluation time to under an hour.