ModelRefs / JailbreakBench Leaderboard — AI Model Scores

JailbreakBench Leaderboard — AI Model Scores

Reproducible jailbreak evaluation across 100 misuse behaviors. Current leaders, methodology, and citation sources for JailbreakBench.

Overview

Reproducible jailbreak evaluation across 100 misuse behaviors.

How it is measured: Attack-success-rate under fixed adversary budget.

How this benchmark is scored

Categorysafety
Maximum score100 % ASR
DirectionLower is better

Primary source: https://jailbreakbench.github.io/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to JailbreakBench Leaderboard — AI Model Scores.