ModelRefs / DROP Methodology — Methodology
DROP Methodology — Methodology
DROP measures discrete reasoning over paragraphs — extraction, counting, comparison, arithmetic.
Overview
What it measures: Reasoning that requires combining multiple facts extracted from a passage.
How it works
- ~96K questions over Wikipedia paragraphs.
- Answers include numbers, dates, and spans.
- Scored on F1 and exact-match.
Strengths
Requires composition of extracted facts
Limitations
- Older benchmark — partially saturated
- F1 normalization noisy
Best use cases
Reading-comprehension baselining
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to DROP Methodology — Methodology.
Frequently asked questions
What does DROP measure?
Reasoning that requires combining multiple facts extracted from a passage.
What are its main limitations?
Older benchmark — partially saturated F1 normalization noisy
When should I use this benchmark?
Reading-comprehension baselining