Chapter 4. Using DuckDB with Polars
Most data scientists and data analysts are familiar with the pandas library. With pandas, you can organize your dataset into Series or DataFrame structures and employ the diverse array of functions provided by the pandas library for data manipulation. However, one of the main complaints about pandas is its slow speed and inefficiencies when dealing with large datasets. This is because pandas was originally designed to work with tabular data that fits in memory. When dealing with large datasets, it becomes slow because it needs to swap data in and out of memory.
To address the inefficiencies of pandas in working with large datasets, there is a competing library—Polars. The first part of this chapter provides an introduction to Polars and how you can work with it (just like with pandas). The second part of this chapter shows how you can query Polars DataFrames using DuckDB.
Introduction to Polars
Polars is a DataFrame library that is completely written in Rust. Polars is designed with the following in mind:
- Speed
Polars leverages Rust, a system programming language known for its performance.
- Parallelism
Polars can take advantage of multicore processors, which provide substantial speed improvements for CPU-bound operations.
- Memory efficiency
Polars uses lazy evaluation, which means an operation is not performed until it is needed. In addition, queries can be chained and optimized before execution, resulting in much more efficient execution.
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access