1. What is a 'jailbreak' in the context of LLMs?
2. What is 'training data bias' and why does it matter for AI safety?
3. How does 'indirect prompt injection' differ from direct prompt injection?
4. What is the purpose of a content moderation classifier as a guardrail?
5. What does 'hallucination' mean in the context of LLM safety?
6. What is the principle of 'minimal footprint' in agentic AI safety?
7. Why is 'human-in-the-loop' oversight especially important for high-stakes AI actions?