Part 1: Introduction and Data Preparation
We begin this book by introducing the foundational concepts necessary to understand and work with large language models (LLMs). In this part, you will explore the critical role of data preparation in building high-quality LLMs. From understanding the significance of design patterns in model development to handling the immense datasets required for training, we guide you through the initial steps of the LLM pipeline. The chapters in this part will help you master data cleaning techniques to improve data quality, data augmentation methods to enhance dataset diversity, and dataset versioning strategies to ensure reproducibility. You will also learn how to efficiently handle large datasets and create well-annotated ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access