Chapter 18. Polars Internals
In this chapter, we dive deep into the inner workings of Polars, uncovering the mechanisms that make it such a powerful and efficient data manipulation package. Throughout the book, you’ve encountered references to this chapter for a more comprehensive understanding of Polars’ functionality.
In this chapter, you’ll learn about:
-
The architecture of Polars
-
Some technologies under the hood
-
Some optimizations that enable Polars’ performance
-
Tools to profile and test your code, ensuring that you can optimize your data processing tasks to the fullest
The instructions to get any files you might need are in Chapter 2. We assume that you have the files in the data subdirectory.
By the end of this chapter, you’ll have a deeper appreciation for the engineering that goes into making Polars fast and efficient, as well as practical knowledge on how to leverage these internals for your own data processing tasks.
Polars’ Architecture
Everything starts with your code. The code you type is in Polars’ domain-specific language (DSL). This DSL allows you to define what should be done in a declarative manner.
The DSL is then transformed into an intermediate representation (IR) that is a more abstract representation of the work that needs to be done.
You can view this IR by calling lf.explain() to get its String representation, or you can visualize it with lf.show_graph().
The IR is represented as a directed acyclic graph (DAG) where each node represents a computation ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access