Chapter 71. The Hidden Cost of Data Input/Output
Lohit VijayaRenu
While data engineers are exposed to libraries and helper functions to read and write data, knowing certain details about the process of reading and writing operations helps you optimize your applications. Understanding and having the ability to configure various options can help with many data-intensive applications, especially at scale. The following are a few hidden details associated with data input/output (I/O).
Data Compression
While everyone agrees that compressed data can save disk space and reduce the cost of network transfer, you can choose from a plethora of compression algorithms for your data. You should always consider the compression speed versus compression ratio while choosing an algorithm. This applies to compression as well as decompression operations. For example, if the data is already heavy, choosing a faster decompression algorithm while sacrificing more resources for slower compression pays off.
Data Format
While most unstructured data is a collection of records, that might not be the best format for certain types of data access. For example, a record that has multiple nested fields, of which only a few are accessed frequently, is best stored in a columnar data format instead of a record format. You should consider using different levels of nesting versus flattening columns for efficient ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access