Notableevaluation benchmarks

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Zhilin Wang, Shaokun Zhang, Yifan Zhang, Hao Zhang, Jin Xu, Binfeng Xu, Jian Hu, Yunheng Zou, Karan Sapra, Andrew Tao, Jan Kautz, Yi Dong

Published
Sep 21, 2026 16:55 UTC

{'Problem': 'Current evaluation methods for Computer-Use Agents (CUAs) lack transparency regarding the reasons for task failures. This paper addresses this gap by proposing OSWorld-Pro, a process-based evaluation framework that provides insights into the specific subgoals that CUAs fail to achieve. The work is presented as a preprint and has not undergone peer review.', 'Method': 'The authors developed a comprehensive task set comprising over 300 tasks, which include more than 2800 subgoals. The evaluation framework is grounded in over 67,000 human annotations, ensuring a robust basis for assessing agent performance. The evaluation method employs human-aligned LLM-Judges to evaluate the fulfillment of subgoals, allowing for a nuanced understanding of agent capabilities and shortcomings.', 'Results': 'The performance of the evaluated agent, Claude Opus 5, was reported at 75.7% on the OSWorld-Pro benchmark, compared to a higher performance of 83.4% on the original OSWorld benchmark. This indicates a notable difference in performance metrics between the two evaluation frameworks.', 'Limitations': 'The authors identify specific failure modes that CUAs encounter, including subgoal-irrelevant actions and click-based mistakes. These limitations highlight areas where CUAs may struggle, but the paper does not discuss additional limitations or potential biases in the evaluation process.', 'Why it matters': 'The introduction of OSWorld-Pro has significant implications for the evaluation of CUAs, as it provides a clearer framework for understanding task failures. This transparency can guide future research in improving agent design and training methodologies, ultimately leading to more effective and reliable CUAs.'}

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI