ModelRefs / Open LLM Leaderboard — AI Glossary
Open LLM Leaderboard — AI Glossary
HuggingFace's public benchmark leaderboard evaluating open-weight models on standardized NLP tasks.
Overview
The Open LLM Leaderboard v2 (2024) evaluates models on MMLU-Pro, GPQA, MuSR, MATH Hard, IFEval, and BBH using the LM Evaluation Harness. Provides a standardized, reproducible comparison of open-weight models. Vulnerable to benchmark contamination as datasets become public; periodically refreshed with harder benchmarks.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Open LLM Leaderboard — AI Glossary.
Frequently asked questions
What is Open LLM Leaderboard?
HuggingFace's public benchmark leaderboard evaluating open-weight models on standardized NLP tasks.
What concepts are related to Open LLM Leaderboard?
Closely related concepts include mmlu, lmsys, bigcodebench.