May 2025
Intermediate to advanced
538 pages
13h 11m
English
In this chapter, we’ll dive into quantization methods that can optimize LLMs for deployment on resource-constrained devices, such as mobile phones, embedded systems, or edge computing environments.
Quantization is a technique that reduces the precision of numerical representations, thus shrinking the model’s size and improving its inference speed without heavily compromising its performance.
Quantization is particularly beneficial in the following scenarios:
Read now
Unlock full access