← Fine-Tuning & Alignment
Alignment
Also known as: model alignment, value alignment
Alignment in the context of language model training refers to the process of ensuring a model's behavior matches intended human values, goals, and policies — covering dimensions such as helpfulness, honesty, and harmlessness. Alignment techniques include instruction tuning, RLHF, DPO, and constitutional AI approaches. Alignment is distinct from capability: a highly capable model that acts deceptively or pursues unintended goals is misaligned even if it performs well on benchmarks.