ModelRefs / Safety Benchmarks — Top AI Models
Safety Benchmarks — Top AI Models
Red-team, jailbreak, toxicity, and refusal-quality evaluations. Quantify resistance to misuse and harmful-output rate.
Overview
Red-team, jailbreak, toxicity, and refusal-quality evaluations.
What this category is for: Quantify resistance to misuse and harmful-output rate.
Benchmarks in this category
- HarmBench — Standardized red-teaming eval covering 510 harmful behaviors.
- JailbreakBench — Reproducible jailbreak evaluation across 100 misuse behaviors.
- ToxiGen — Implicit toxicity classification across 13 demographic groups.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Safety Benchmarks — Top AI Models.