Frontier Engineering
AI Safety, Ethics & Risk

Bias

Bias in AI systems refers to systematic errors in model outputs that reflect or amplify unequal treatment of demographic groups, topics, or viewpoints. It can originate in skewed training data, the choice of optimization objective, or the demographic composition of annotators who provide preference labels. Bias manifests as lower accuracy for underrepresented groups, stereotypical associations in language model outputs, or disproportionate toxicity predictions directed at particular communities.