ModelRefs / MT-Bench — AI Glossary
MT-Bench — AI Glossary
A multi-turn chat evaluation benchmark using GPT-4 as judge to rate models on 80 two-turn questions across 8 domains. Scores range 1–10.
Overview
MT-Bench (Zheng et al. 2023, LMSYS) tests instruction following in multi-turn conversations across writing, roleplay, extraction, reasoning, math, coding, STEM, and humanities. GPT-4 scores as a judge correlate well with human preferences. Used alongside Chatbot Arena to rank chat models. Scores range 1–10.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MT-Bench — AI Glossary.
Frequently asked questions
What is MT-Bench?
A multi-turn chat evaluation benchmark using GPT-4 as judge to rate models on 80 two-turn questions across 8 domains.
What concepts are related to MT-Bench?
Closely related concepts include alpacaeval, chatbot arena, lmsys.