ModelRefs / WinoGrande — AI Glossary
WinoGrande — AI Glossary
A large-scale Winograd schema benchmark testing commonsense pronoun disambiguation at scale. Harder than the original Winograd Schema Challenge.
Overview
WinoGrande (Sakaguchi et al. 2019) contains 44,000 pronoun resolution problems adversarially filtered to remove annotator artifacts. Each problem requires physical, social, or causal commonsense to resolve the correct antecedent. Harder than the original Winograd Schema Challenge. GPT-4 scores ~91%; 7B models typically 70–80%.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to WinoGrande — AI Glossary.
Frequently asked questions
What is WinoGrande?
A large-scale Winograd schema benchmark testing commonsense pronoun disambiguation at scale.
What concepts are related to WinoGrande?
Closely related concepts include hellaswag, arc challenge, commonsense reasoning.