February 2019
Beginner to intermediate
544 pages
14h 36m
English
Files are important parts of a data pipeline as data is stored in files. We use different file formats for data management, configuration management, or for other uses. For example, the common files we have been using are XML, JSON, and CSV. Users may require a more lightweight data format and web service for Hadoop. The CPU and memory needed for processing the extra bytes, transporting over the network, and storing using storage are always limited and therefore it is important to choose the right file format and compression techniques. Let's look into how we do it.
Read now
Unlock full access