ModelRefs / HellaSwag Leaderboard — AI Model Scores

HellaSwag Leaderboard — AI Model Scores

Commonsense NLI: select the most plausible 4-way continuation of an ActivityNet/WikiHow paragraph. Current leaders, methodology, and citation sources for HellaSwag.

Overview

Commonsense NLI: select the most plausible 4-way continuation of an ActivityNet/WikiHow paragraph.

How it is measured: 10-ending accuracy over the validation split; adversarially filtered to be hard for pre-2019 models.

How this benchmark is scored

Categoryopen-source
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://rowanzellers.com/hellaswag/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to HellaSwag Leaderboard — AI Model Scores.