December 2016
Intermediate to advanced
848 pages
26h 22m
English
This chapter covers the following:
Safeguarding HDFS data using trash and HDFS snapshots
Ensuring data integrity with file system checks (fsck command)
File-based formats supported by Hadoop
Choosing the optimal file format
The Hadoop small files problem and merging files
Using Hadoop archives ...
Read now
Unlock full access