ModelRefs / BIG-Bench Hard (BBH) — AI Glossary
BIG-Bench Hard (BBH) — AI Glossary
A 23-task subset of BIG-Bench comprising tasks on which prior language models performed below human level. Chain-of-thought unlocks large gains.
Overview
BBH (Suzgun et al. 2022) filters BIG-Bench tasks where GPT-3 averaged <65% to identify genuinely hard challenges: causal reasoning, logical deduction, algorithmic tasks, dyck languages. Chain-of-thought unlocks large gains. GPT-4 scores ~85%; a useful robustness check alongside MMLU.
Reference details
| Topic | evaluation |
|---|---|
| Also known as | BIG-Bench Hard |
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BIG-Bench Hard (BBH) — AI Glossary.
Frequently asked questions
What is BIG-Bench Hard (BBH)?
A 23-task subset of BIG-Bench comprising tasks on which prior language models performed below human level.
Is BIG-Bench Hard (BBH) the same as BIG-Bench Hard?
Yes — BIG-Bench Hard are common aliases for BIG-Bench Hard (BBH).
What concepts are related to BIG-Bench Hard (BBH)?
Closely related concepts include mmlu, arc challenge, chain of thought.