ModelRefs / AlpacaEval 2 LC Leaderboard — AI Model Scores
AlpacaEval 2 LC Leaderboard — AI Model Scores
Length-controlled win-rate vs GPT-4 Turbo on 805 instruction-following prompts. Current leaders, methodology, and citation sources for AlpacaEval 2 LC.
Overview
Length-controlled win-rate vs GPT-4 Turbo on 805 instruction-following prompts.
How it is measured: LC win-rate: GPT-4 Turbo as judge; length-controlled to reduce verbosity bias.
How this benchmark is scored
| Category | open-source |
|---|---|
| Maximum score | 100 % LC win |
| Direction | Higher is better |
Primary source: https://tatsu-lab.github.io/alpaca_eval/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AlpacaEval 2 LC Leaderboard — AI Model Scores.