4
Foundations of AI Defense
AI attack surfaces are broad enough that it is easy to lose track of what to defend first. Natural language interfaces create an effectively unbounded input space, and long-lived state and memory introduce failure modes that look nothing like traditional web app bugs. Tool integrations push model decisions into code execution, storage systems, and remote services, so a single prompt can now reach places that used to be several hops away from any user input. Controls built for static APIs do not line up neatly with this behavior.
Defenders have not been starting from scratch, though. Those incidents covered in Chapter 3 forced vendors and security teams to try many ideas in production: Sydney's prompt injection issues ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access