ModelRefs / HarmBench Leaderboard — AI Model Scores
HarmBench Leaderboard — AI Model Scores
Standardized red-teaming eval covering 510 harmful behaviors. Current leaders, methodology, and citation sources for HarmBench.
Overview
Standardized red-teaming eval covering 510 harmful behaviors.
How it is measured: Attack success rate against classifier (lower better).
How this benchmark is scored
| Category | safety |
|---|---|
| Maximum score | 100 % ASR |
| Direction | Lower is better |
Primary source: https://www.harmbench.org/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to HarmBench Leaderboard — AI Model Scores.