Frontier Engineering
AI Safety, Ethics & Risk

Jailbreak

A jailbreak is a user-crafted prompt designed to circumvent a language model's safety training and elicit behavior the model was explicitly trained to refuse — such as providing instructions for harmful activities. Jailbreaks exploit the tension between helpfulness and safety by using roleplay framing, hypothetical scenarios, or obfuscation to fool safety classifiers or instruction-following logic. Jailbreak robustness is evaluated through adversarial red-teaming and drives iterative safety training improvements.