May 2026
Intermediate
376 pages
9h 52m
English
In chapter 5, you learned the core ONNX concepts and capabilities. I only briefly mentioned model quantization, but it’s crucial for LLM inference performance, so it deserves its own chapter. That’s our focus here.
Previous chapters have mentioned several numeric precision formats used for LLM training and inference. It’s time to look at them more closely and to introduce additional formats used for model quantization.
In traditional scientific computing, 64-bit floating point (double ...
Read now
Unlock full access