CHAPTER 7Optimizing Performance for Foundational Models
Generative AI models, known for their immense scale and complexity, have unlocked unprecedented capabilities in areas like natural language processing and generative art. While these models drive advanced end-to-end solutions, they also present challenges in computational efficiency and resource utilization, making optimization essential.
This chapter delves into the multifaceted endeavor of optimizing performance for foundational models, breaking the topic into key areas that impact model execution and efficiency. We begin by exploring the challenges of compute and memory with large language models (LLMs), including memory overhead, compute constraints, and inference optimization techniques to ensure scalable and cost-effective execution.
To refine model performance, we discuss evaluation and model refinement strategies, covering reinforcement learning (RL) for improving model responses, automatic and human-driven model evaluations, and specific model evaluation tasks such as text generation, summarization, Q&A, and classification. Additionally, we discuss the importance of crafting appropriate prompt datasets, differentiating between built-in and custom prompt datasets, and analyzing evaluation results through automated reports and human review processes.
Optimization strategies extend beyond model evaluation, requiring efficient distributed computing approaches. We examine data parallelism, model parallelism, and hybrid ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access