AI security is the defense of machine learning systems against attacks on the models themselves: adversarial attacks on inputs, poisoning of training data, prompt injection into LLM applications, and the jailbreaks that follow. The field gained shared vocabulary when OWASP put prompt injection at the top of its 2025 Top 10 for LLM applications, ahead of sensitive information disclosure and poisoning, and published the mitigations that only partially work . MITRE ATLAS maintains the adversary-side map, a living knowledge base of tactics and techniques built from real-world observations and red team demonstrations . Employers now need engineers who read both and can turn them into architecture, and the population that can is thin.
Challenges in AI Security Recruiting
Adversarial attacks became an operational risk with a taxonomy
NIST's AI 100-2e2025 taxonomy turned adversarial attacks from a paper phenomenon into a classified threat set: evasion, poisoning and privacy attacks for predictive systems; poisoning, direct and indirect prompt injection for generative ones; each with named mitigations and documented limits . The report separates attacks by the property they violate, availability, integrity, privacy, and by where in the lifecycle they land, which is the frame an incident response team now expects its engineers to speak fluently . MITRE ATLAS mirrors it for defenders, with technique IDs that security teams already know how to consume from ATT&CK . The taxonomy also documents where widely used mitigations break down, which matters because certified defense algorithms, the branch of the literature with provable robustness guarantees, remain a research track rather than a deployment default . For recruiters the practical split runs between the two populations: adversarial ML researchers who publish new attacks, and defenders who turn taxonomies into controls. Teams that hire pure ML and expect the security to appear later discover the gap at the first audit.
Prompt injection defense has no complete fix, only layered mitigations
Prompt injection defense is unique among security problems because the instructions and the data share one channel. OWASP states plainly that no complete prevention exists: the mitigations are constraining the model's role, input and output filtering, least-privilege tool access and human approval for high-risk actions . It also catalogs the attack shapes defenders must watch: direct injections into the prompt itself, indirect injections smuggled through retrieved content, code injection through model-connected tools, and the adversarial suffix strings that bypass safety measures . That changes the job description. A defender here is not building a patch; they are composing layers and keeping them ranked against new attacks. Candidates who have run this loop talk about which layer caught which attack and where the chain still leaks. Candidates who have only read the Top 10 talk about the Top 10.
Data poisoning prevention starts at the supply chain
Data poisoning prevention is mostly a supply chain discipline. Training data, checkpoints and embeddings arrive from partners, public hubs and scrapes, and SAFE-AI, MITRE's framework for securing AI-enabled systems, points at unclear provenance and third-party models as core exposure, mapping each ATLAS threat to NIST SP 800-53 controls across environment, platform, model and data . The adversarial ML threat matrix built with Microsoft records the field's case law: poisoning incidents against production models, model replication attempts, models corrupted after deployment . That record matters for hiring because it shows the attacks are not hypothetical; they have happened to shipped systems, and the people who defended against them learned on the job. The engineering work is provenance: dataset lineage, model cards enforced at ingestion, canaries, periodic re-evaluation for backdoors. SAFE-AI's structure shows how deep the role cuts, since it evaluates controls per system element, environment, platform, model and data, rather than against a single security perimeter . A security engineer who has never seen a training pipeline cannot do this job; a data engineer who has never seen a threat model can only do half of it. The role sits at that seam.
Jailbreak detection runs as a standing evaluation workload
Jailbreak detection is an operations commitment, not a feature. OWASP distinguishes jailbreaking from ordinary prompt injection: the attacker drives the model to disregard its safety behavior entirely, and staying ahead requires ongoing updates to training and safety mechanisms rather than a one-time filter . ATLAS records the technique family as its own entries, which is how security operations track it in the wild . Teams that are serious run jailbreak corpora against every release, track refusal rates and over-refusal, and feed findings back into fine-tuning. The engineers who have held that loop are a different population from the ones who ran a benchmark once, and the interview should confirm which one is in the room. Interview questions should ask which corpus, which release cadence, and what changed after the last run.
Model inversion protection imports privacy engineering into AI security
Model inversion protection sits at the border between AI security and privacy engineering. NIST's taxonomy groups reconstruction and membership inference with model extraction under privacy attacks on predictive systems, alongside property inference variants, and its glossary exists because the field could not agree on terms . Defense work here is statistical: training with bounds that limit what gradients reveal, output checks against reconstruction, access controls on inference APIs. The grading question is always the same, how much of a training record an attacker can recover from the deployed model and at what query cost. Employers reach into two pools, security and privacy, and find almost no one who has done this work on production models. The brief usually ends up hiring for adjacent skills and building the specialty in-house, which means the first person through the door is effectively defining the function.
Secure AI architecture claims fail at the threat-model question
Secure AI architecture is easy to claim because the phrase has no standard meaning. The verification step is the threat model. Ask what the adversary knows, which stage of the lifecycle they attack, which ATLAS techniques the design counters, and which SP 800-53 controls the environment maps to . Ask how the team tests: NIST's taxonomy is explicit that common mitigations carry known limits, so a candidate should be able to say where their defenses stop working . Ask what a robustness testing program looks like in practice: which perturbations, which corpora, which regression gate, and what happens to a finding when it lands . People who have built secure AI systems answer with a named adversary and a control map. People who have read about them answer with vocabulary, at length. The cost of hiring the second kind is a system that passes its initial review and fails its first real attack, usually after it has shipped.
References
- OWASP Top 10 for LLM Applications 2025 — OWASP Foundation. (accessed 2026-09-28)
- MITRE ATLAS: Adversarial Threat Landscape for Artificial-Intelligence Systems — MITRE. (accessed 2026-09-28)
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025) — National Institute of Standards and Technology (NIST). (accessed 2026-09-28)
- SAFE-AI: A Framework for Securing AI-Enabled Systems — MITRE. (accessed 2026-09-28)
- Adversarial ML Threat Matrix — MITRE and Microsoft. (accessed 2026-09-28)
