← AI Safety, Ethics & Risk
Adversarial Input
An adversarial input is a carefully crafted example designed to cause a model to produce an incorrect, unsafe, or unintended output — often while appearing innocuous to a human reviewer. In text models, adversarial inputs include character substitutions, homoglyphs, unusual Unicode sequences, or semantics-preserving paraphrases that fool safety classifiers. Adversarial robustness testing reveals brittleness in both safety systems and core model capabilities before deployment.