Chapter 12. Threat Modeling for AI Agents
When you’re deploying AI agents, you need to build a layered, asset-centric threat model for systems that think, act, and persist state. You might already know the basics from general information security standards. However, traditional systems don’t decide—agents do—which is why you need to extend classical threat modeling to nondeterministic, agentic systems. Agent security is often approached as a collection of isolated vulnerabilities such as prompt injection or tool misuse. However, in agentic systems, risks rarely remain confined to a single component. A failure at the model layer can propagate through orchestration, trigger unintended tool execution, and persist in memory, affecting future decisions. To reason about these systems, it’s helpful to use a layered threat model that maps assets, system layers, and threat propagation paths. This allows you to treat security not as a checklist, but as a system design problem under constraints.
To build this understanding, you can draw from established security thinking without applying it one-to-one. Traditional approaches emphasize identifying assets, classifying their protection levels, and systematically analyzing threats across system layers. Modern application security frameworks add concrete attack patterns such as injection, data exfiltration, or misuse of execution interfaces. And red teaming provides a way to validate whether these risks can actually be exploited in your system. ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access