ModelRefs / HarmBench Leaderboard — AI Model Scores

HarmBench Leaderboard — AI Model Scores

Standardized red-teaming eval covering 510 harmful behaviors. Current leaders, methodology, and citation sources for HarmBench.

Overview

Standardized red-teaming eval covering 510 harmful behaviors.

How it is measured: Attack success rate against classifier (lower better).

How this benchmark is scored

Categorysafety
Maximum score100 % ASR
DirectionLower is better

Primary source: https://www.harmbench.org/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to HarmBench Leaderboard — AI Model Scores.