Frontier Engineering

Engineering

AI Security Engineer / Red Teamer

Break AI systems on purpose, then design the controls that make those attacks fail.

Role overview

AI security engineers own the question nobody else on the team is paid to ask: what happens when the text reaching the model is written by someone who wants to hurt you. That covers prompt injection arriving through retrieved documents and browsed pages, agents tricked into calling tools with a user's privileges, data leaving through rendered output, poisoned fine-tuning corpora, and model artifacts pulled from public hubs with nobody checking what they deserialise.

The work splits into two halves that reinforce each other. One half is offensive: scoping and running red-team exercises, probing refusal boundaries systematically, and writing findings that engineering will actually fix. The other half is defensive architecture — trust boundaries, tool authorisation, sandboxed execution, tenancy isolation, and adversarial test suites wired into CI so regressions fail a build instead of a customer.

Interviews for this role probe whether you reason about AI risk structurally or reach for a classifier every time. Strong candidates talk about permission models and blast radius, quantify attack success rate rather than collecting anecdotes, and are candid about which defences are genuinely weak.

Skills and stack

Attack surface

  • Direct and indirect prompt injection
  • Tool and agent abuse, confused-deputy paths
  • Data exfiltration through rendered output and outbound tools
  • Retrieval index poisoning and persistent agent memory
  • Jailbreak and refusal-boundary probing
  • Economic abuse and denial-of-wallet patterns

Defensive architecture

  • Trust boundaries and untrusted-content framing
  • Tool authorisation scoped to the calling user
  • Sandboxed code execution and egress control
  • Multi-tenant isolation across stores, caches, and corpora
  • Human confirmation for irreversible actions
  • Defence in depth and blast-radius reduction

Supply chain and training

  • Checkpoint provenance, pinning, and safe serialisation formats
  • Training-data poisoning and backdoor trigger hunting
  • Dataset contribution limits and provenance metadata
  • Inference-stack dependency review

Testing and measurement

  • Adversarial benchmark suites and attack success rate
  • Canary secrets and egress detection
  • False-block rate on real traffic
  • Per-model-version regression gates

Programme and communication

  • Threat modelling an LLM feature
  • Rules of engagement and red-team scoping
  • Findings reports with reproduction rates and severity
  • Incident response for agent actions
  • Working with product teams without becoming a launch blocker

Interview questions

Expand a question to read a model answer. Filter by focus area or seniority to rehearse the rounds you are actually facing.

Focus
Level

Showing 20 of 20 questions

Rehearse it out loud.

Reading model answers is not the same as saying one under pressure. Book a 30-minute 1:1 and run a mock AI Security Engineer / Red Teamer interview — scored, with the gaps named while they are still cheap to fix.