18
Adversarial Robustness
Adversarial attacks on LLMs are designed to manipulate the model’s output by making small, often imperceptible changes to the input. These attacks can expose vulnerabilities in LLMs and potentially lead to security risks or unintended behaviors in real-world applications.
In this chapter, we’ll discover techniques for creating and defending against adversarial examples in LLMs. Adversarial examples are carefully crafted inputs designed to intentionally mislead the model into producing incorrect or unexpected outputs. You’ll learn about textual adversarial attacks, methods to generate these examples, and techniques to make your models more robust. We’ll also cover evaluation methods and discuss the real-world implications ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access