← AI Safety, Ethics & Risk
Jailbreak
A jailbreak is a user-crafted prompt designed to circumvent a language model's safety training and elicit behavior the model was explicitly trained to refuse — such as providing instructions for harmful activities. Jailbreaks exploit the tension between helpfulness and safety by using roleplay framing, hypothetical scenarios, or obfuscation to fool safety classifiers or instruction-following logic. Jailbreak robustness is evaluated through adversarial red-teaming and drives iterative safety training improvements.