ModelRefs / TruthfulQA Methodology — Methodology

TruthfulQA Methodology — Methodology

TruthfulQA measures whether models produce truthful answers to questions where humans hold common misconceptions.

Overview

What it measures: Resistance to imitating human falsehoods learned from training data.

How it works

  • 817 questions across 38 categories.
  • Each question has known true and false reference answers.
  • Scored by an LLM judge (GPT-judge) or via multiple-choice variants.

Strengths

Probes a specific failure mode (misconception imitation)

Limitations

  • Judge-model bias
  • Small set
  • Categories overlap

Best use cases

Hallucination-resistance evaluation

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to TruthfulQA Methodology — Methodology.

Frequently asked questions

What does TruthfulQA measure?

Resistance to imitating human falsehoods learned from training data.

What are its main limitations?

Judge-model bias Small set Categories overlap

When should I use this benchmark?

Hallucination-resistance evaluation