ModelRefs / TruthfulQA Methodology — Methodology
TruthfulQA Methodology — Methodology
TruthfulQA measures whether models produce truthful answers to questions where humans hold common misconceptions.
Overview
What it measures: Resistance to imitating human falsehoods learned from training data.
How it works
- 817 questions across 38 categories.
- Each question has known true and false reference answers.
- Scored by an LLM judge (GPT-judge) or via multiple-choice variants.
Strengths
Probes a specific failure mode (misconception imitation)
Limitations
- Judge-model bias
- Small set
- Categories overlap
Best use cases
Hallucination-resistance evaluation
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to TruthfulQA Methodology — Methodology.
Frequently asked questions
What does TruthfulQA measure?
Resistance to imitating human falsehoods learned from training data.
What are its main limitations?
Judge-model bias Small set Categories overlap
When should I use this benchmark?
Hallucination-resistance evaluation