ModelRefs / OSWorld — AI Glossary

OSWorld — AI Glossary

A benchmark evaluating computer-using agents on real desktop tasks across Ubuntu, macOS, and Windows. GPT-4V at launch: ~13% success; humans: 72%.

Overview

OSWorld (Xie et al. 2024) presents 369 real computer tasks (LibreOffice editing, Chrome navigation, file management, VsCode use) with functional verification. Agents observe screenshots and generate mouse/keyboard actions. GPT-4V at launch: ~13% success; humans: 72%. Defines the frontier for GUI/CUA evaluation.

Reference details

Topicevaluation
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to OSWorld — AI Glossary.

Frequently asked questions

What is OSWorld?

A benchmark evaluating computer-using agents on real desktop tasks across Ubuntu, macOS, and Windows.

What concepts are related to OSWorld?

Closely related concepts include computer using agent, gaia, web agent.