ModelRefs / Alignment Tax — AI Glossary

Alignment Tax — AI Glossary

The degradation in raw capability that may result from applying alignment and safety training procedures to a base language model.

Overview

Alignment tax refers to performance drops on benchmarks after RLHF or Constitutional AI—the model may be safer but less capable on certain tasks. Empirically, the tax is small for frontier models: InstructGPT retained near-GPT-3 quality while dramatically improving helpfulness. Ongoing research aims to reduce or eliminate the trade-off.

Reference details

Topicsafety
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Alignment Tax — AI Glossary.

Frequently asked questions

What is Alignment Tax?

The degradation in raw capability that may result from applying alignment and safety training procedures to a base language model.

What concepts are related to Alignment Tax?

Closely related concepts include safety training, alignment, rlhf.