ModelRefs / CRUXEval Leaderboard — AI Model Scores
CRUXEval Leaderboard — AI Model Scores
Code reasoning over input prediction and output prediction. Current leaders, methodology, and citation sources for CRUXEval.
Overview
Code reasoning over input prediction and output prediction.
How it is measured: Pass@1 averaged across input and output tasks.
How this benchmark is scored
| Category | coding |
|---|---|
| Maximum score | 100 pass@1 |
| Direction | Higher is better |
Primary source: https://crux-eval.github.io/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to CRUXEval Leaderboard — AI Model Scores.