ModelRefs / SimpleQA Leaderboard — AI Model Scores
SimpleQA Leaderboard — AI Model Scores
Factuality benchmark of short single-fact questions. Current leaders, methodology, and citation sources for SimpleQA.
Overview
Factuality benchmark of short single-fact questions.
How it is measured: Graded F1 (correct − incorrect rate).
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % correct |
| Direction | Higher is better |
Primary source: https://openai.com/index/introducing-simpleqa/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SimpleQA Leaderboard — AI Model Scores.