ModelRefs / Alignment Tax — AI Glossary
Alignment Tax — AI Glossary
The degradation in raw capability that may result from applying alignment and safety training procedures to a base language model.
Overview
Alignment tax refers to performance drops on benchmarks after RLHF or Constitutional AI—the model may be safer but less capable on certain tasks. Empirically, the tax is small for frontier models: InstructGPT retained near-GPT-3 quality while dramatically improving helpfulness. Ongoing research aims to reduce or eliminate the trade-off.
Reference details
| Topic | safety |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Alignment Tax — AI Glossary.
Frequently asked questions
What is Alignment Tax?
The degradation in raw capability that may result from applying alignment and safety training procedures to a base language model.
What concepts are related to Alignment Tax?
Closely related concepts include safety training, alignment, rlhf.