← AI Safety, Ethics & Risk
AI Alignment
AI alignment is the research and engineering discipline concerned with ensuring that AI systems pursue goals that are genuinely beneficial to humans. A misaligned model may optimize for a measurable proxy objective — maximizing engagement, minimizing loss — while producing outcomes that deviate from human intentions. Alignment work spans specification (defining the right objectives), training (instilling those objectives), and verification (confirming that deployed systems behave as intended across novel situations).