ModelRefs / MuSR Leaderboard — AI Model Scores
MuSR Leaderboard — AI Model Scores
Multistep soft reasoning over narrative scenarios. Current leaders, methodology, and citation sources for MuSR.
Overview
Multistep soft reasoning over narrative scenarios.
How it is measured: Zero-shot accuracy on murder mystery / team allocation / object placements.
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://arxiv.org/abs/2310.16049
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MuSR Leaderboard — AI Model Scores.