Blog

AI Agent Glossary

OSWorld

The reference benchmark for computer-use agents: 369 real tasks on real operating systems.

OSWorld runs agents inside actual Ubuntu, Windows and macOS VMs and scores them with 134 execution-based evaluators (checking files, application state) rather than text matching. Humans complete 72.36% of tasks. OSWorld-Verified (July 2025) fixed community-reported task issues and cut evaluation time to under an hour.

Source: OSWorld — Xie et al., 2024

See the numbers behind this term: AI Agent Statistics 2026.