Frontier Engineering
AI Safety, Ethics & Risk

Toxicity

Toxicity refers to model outputs that are harmful, offensive, or hurtful — including hate speech, personal attacks, profanity, or content that demeans individuals based on protected characteristics. Language models trained on internet text learn to produce toxic outputs because such text is prevalent in pretraining corpora. Measuring toxicity requires probabilistic classifiers that are themselves imperfect and can carry their own biases, requiring careful calibration and human review in content moderation pipelines.