ModelRefs / BBH (Big-Bench Hard) Methodology — Methodology
BBH (Big-Bench Hard) Methodology — Methodology
BBH is the subset of BIG-Bench tasks where prior LLMs underperformed humans.
Overview
What it measures: A curated hard subset of diverse reasoning tasks.
How it works
- 23 tasks from BIG-Bench.
- Mix of algorithmic, logical, and commonsense reasoning.
- Typically evaluated with chain-of-thought.
Strengths
- Task diversity
- Discriminative across model families
Limitations
Heterogeneous scoring complicates aggregation
Best use cases
Reasoning-quality differentiation
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BBH (Big-Bench Hard) Methodology — Methodology.
Frequently asked questions
What does BBH (Big-Bench Hard) measure?
A curated hard subset of diverse reasoning tasks.
What are its main limitations?
Heterogeneous scoring complicates aggregation
When should I use this benchmark?
Reasoning-quality differentiation