← AI Safety, Ethics & Risk
Content Moderation
Content moderation in AI systems is the process of reviewing and filtering model inputs and outputs to prevent the generation or amplification of harmful content — including hate speech, graphic violence, sexual content, misinformation, and self-harm material. Automated moderation uses classifier models tuned for specific harm categories; human review adds a quality check for edge cases. Effective moderation pipelines combine multiple classifiers with configurable sensitivity thresholds and escalation paths to human reviewers.