ModelRefs / HellaSwag Leaderboard — AI Model Scores
HellaSwag Leaderboard — AI Model Scores
Commonsense NLI: select the most plausible 4-way continuation of an ActivityNet/WikiHow paragraph. Current leaders, methodology, and citation sources for HellaSwag.
Overview
Commonsense NLI: select the most plausible 4-way continuation of an ActivityNet/WikiHow paragraph.
How it is measured: 10-ending accuracy over the validation split; adversarially filtered to be hard for pre-2019 models.
How this benchmark is scored
| Category | open-source |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://rowanzellers.com/hellaswag/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to HellaSwag Leaderboard — AI Model Scores.