The alignment tax is the performance cost a model pays for safety training — RLHF and refusal tuning can reduce raw capability on benchmarks while making the model safer and more useful in practice. Minimizing this tax is an active area of research, as the goal is to make models both safe and capable without significant trade-offs.