May 2026
Intermediate
376 pages
9h 52m
English
Chapter 5 introduced the ONNX format and ONNX Runtime capabilities, first in general and then for LLMs. Chapter 6 covered several approaches, including 8-bit quantization through the ONNX API. This chapter will explore other ONNX features, such as profiling the performance of LLMs ported to ONNX format and tools you can use to extract useful insights from the profiling data.
The ONNX Runtime (ORT) delivers high performance for running machine learning (ML) and deep learning (DL) models across a wide range of hardware. Still, to meet specific key performance ...
Read now
Unlock full access