ModelRefs / BIG-Bench Hard (BBH) — AI Glossary

BIG-Bench Hard (BBH) — AI Glossary

A 23-task subset of BIG-Bench comprising tasks on which prior language models performed below human level. Chain-of-thought unlocks large gains.

Overview

BBH (Suzgun et al. 2022) filters BIG-Bench tasks where GPT-3 averaged <65% to identify genuinely hard challenges: causal reasoning, logical deduction, algorithmic tasks, dyck languages. Chain-of-thought unlocks large gains. GPT-4 scores ~85%; a useful robustness check alongside MMLU.

Reference details

Topicevaluation
Also known asBIG-Bench Hard
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BIG-Bench Hard (BBH) — AI Glossary.

Frequently asked questions

What is BIG-Bench Hard (BBH)?

A 23-task subset of BIG-Bench comprising tasks on which prior language models performed below human level.

Is BIG-Bench Hard (BBH) the same as BIG-Bench Hard?

Yes — BIG-Bench Hard are common aliases for BIG-Bench Hard (BBH).

What concepts are related to BIG-Bench Hard (BBH)?

Closely related concepts include mmlu, arc challenge, chain of thought.