Skip to Content
DuckDB: Up and Running
book

DuckDB: Up and Running

by Wei-Meng Lee
December 2024
Intermediate to advanced
308 pages
6h 43m
English
O'Reilly Media, Inc.
Content preview from DuckDB: Up and Running

Chapter 7. Using DuckDB with JupySQL

Traditionally, data scientists use Jupyter Notebook to pull data from database servers or from external datasets (such as CSV, JSON files, etc.) and store it into pandas DataFrames (see Figure 7-1).

Traditional way of querying data as pandas DataFrames and then using them for data visualization
Figure 7-1. Traditional way of querying data as pandas DataFrames and then using them for data visualization

They then use the DataFrames for visualization purposes. This approach has a couple of drawbacks:

  • Querying a database server may degrade the performance of the database server, which may not be optimized for analytical workloads.

  • Loading the data into DataFrames takes up precious resources, including memory and compute. For example, if the intention is to visualize certain aspects of the dataset, you need to load the entire dataset into memory before you can perform visualization on it.

  • Plotting visualizations using Matplotlib also uses a significant amount of memory. Behind the scenes, Matplotlib maintains various objects such as figures, axes, lines, text, and other graphical elements in memory. Each of these elements consumes resources as they are created and rendered. Additionally, Matplotlib handles data arrays used for plotting and temporarily stores them in memory for processing. If you’re creating multiple plots or figures, each figure and its associated data remain in memory until explicitly closed or cleared, leading to increased ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

DuckDB in Action

DuckDB in Action

Mark Needham, Michael Hunger, Michael Simons
FastAPI

FastAPI

Bill Lubanovic
Kubernetes: Up and Running, 3rd Edition

Kubernetes: Up and Running, 3rd Edition

Brendan Burns, Joe Beda, Kelsey Hightower, Lachlan Evenson
Docker: Up & Running, 3rd Edition

Docker: Up & Running, 3rd Edition

Sean P. Kane, Karl Matthias

Publisher Resources

ISBN: 9781098159689Errata Page