ModelRefs / BBH (Big-Bench Hard) Methodology — Methodology

BBH (Big-Bench Hard) Methodology — Methodology

BBH is the subset of BIG-Bench tasks where prior LLMs underperformed humans.

Overview

What it measures: A curated hard subset of diverse reasoning tasks.

How it works

  • 23 tasks from BIG-Bench.
  • Mix of algorithmic, logical, and commonsense reasoning.
  • Typically evaluated with chain-of-thought.

Strengths

  • Task diversity
  • Discriminative across model families

Limitations

Heterogeneous scoring complicates aggregation

Best use cases

Reasoning-quality differentiation

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BBH (Big-Bench Hard) Methodology — Methodology.

Frequently asked questions

What does BBH (Big-Bench Hard) measure?

A curated hard subset of diverse reasoning tasks.

What are its main limitations?

Heterogeneous scoring complicates aggregation

When should I use this benchmark?

Reasoning-quality differentiation