ModelRefs / LongFact Concepts Leaderboard — AI Model Scores

LongFact Concepts Leaderboard — AI Model Scores

Long-form factuality benchmark measuring unsupported claims in open-ended concept answers. Current leaders, methodology, and citation sources for LongFact Concepts.

Overview

Long-form factuality benchmark measuring unsupported claims in open-ended concept answers.

How it is measured: Claim-level hallucination rate using a browsing-enabled model grader; lower is better.

How this benchmark is scored

Categoryreasoning
Maximum score100 % hallucination rate
DirectionLower is better

Primary source: https://arxiv.org/abs/2403.18802

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to LongFact Concepts Leaderboard — AI Model Scores.