ModelRefs / Open LLM Leaderboard — AI Glossary

Open LLM Leaderboard — AI Glossary

HuggingFace's public benchmark leaderboard evaluating open-weight models on standardized NLP tasks.

Overview

The Open LLM Leaderboard v2 (2024) evaluates models on MMLU-Pro, GPQA, MuSR, MATH Hard, IFEval, and BBH using the LM Evaluation Harness. Provides a standardized, reproducible comparison of open-weight models. Vulnerable to benchmark contamination as datasets become public; periodically refreshed with harder benchmarks.

Reference details

Topicevaluation
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Open LLM Leaderboard — AI Glossary.

Frequently asked questions

What is Open LLM Leaderboard?

HuggingFace's public benchmark leaderboard evaluating open-weight models on standardized NLP tasks.

What concepts are related to Open LLM Leaderboard?

Closely related concepts include mmlu, lmsys, bigcodebench.