ModelRefs / WinoGrande Leaderboard — AI Model Scores

WinoGrande Leaderboard — AI Model Scores

Large-scale Winograd-schema commonsense pronoun resolution (WinoGrande XL). Current leaders, methodology, and citation sources for WinoGrande.

Overview

Large-scale Winograd-schema commonsense pronoun resolution (WinoGrande XL).

How it is measured: Binary pronoun selection accuracy on the xl split (1267 questions); debiased with AFLITE.

How this benchmark is scored

Categoryopen-source
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://winogrande.allenai.org/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to WinoGrande Leaderboard — AI Model Scores.