ModelRefs / OSWorld Leaderboard — AI Model Scores

OSWorld Leaderboard — AI Model Scores

Computer-use agents across real OS apps (Ubuntu/Windows/macOS). Current leaders, methodology, and citation sources for OSWorld.

Overview

Computer-use agents across real OS apps (Ubuntu/Windows/macOS).

How it is measured: 369 tasks; success rate with screenshot loop.

How this benchmark is scored

Categoryagents
Maximum score100 % success
DirectionHigher is better

Primary source: https://os-world.github.io/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to OSWorld Leaderboard — AI Model Scores.