Part 2
Attacks and Agent Behavior
The second part of the book focuses on AI-specific attacks and the behaviors that make them work. The point is not to collect scary examples. The point is to understand the mechanics well enough to test your own systems and avoid false confidence.
This part begins by decomposing an AI system into its working parts. From there, it examines prompt injection and jailbreaks, memory contamination, RAG attacks, agent architecture, agent exploitation, model integrity, training data compromise, and AI red teaming. These chapters show how attackers move through language, context, state, retrieval, tools, delegated authority, and model supply chains.
Chapter 5 maps the system. Chapter 6 shows how instructions become attack ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access