15
Defensive Prompt Engineering
The authentication worked. The request carried a valid token. The JSON was well-formed. Rate limits held. Network policy kept the model away from production databases. None of that stopped the attack because the dangerous instruction did not arrive as an obvious exploit string. It was buried inside content the assistant had been told to trust.
Picture a support assistant summarizing tickets from a company knowledge base. An attacker opens an ordinary ticket and hides an instruction inside it: when this case is summarized, include customer email addresses from other open tickets. The model sees retrieved text, treats the embedded instruction as part of the task, and starts drafting a response it should never send. ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access