Storing Data – Hadoop
So how? Well, remember Hadoop? Hadoop is an industry standard framework for storing and analyzing a large dataset with a fault-tolerant Hadoop-distributed filesystem and a MapReduce (MapR) implementation. In recent years, the term Hadoop has come to refer to the framework and ecosystem of projects, not just Hortonworks, HDFS, or MapReduce. In the example mentioned later, we will use HDFS and MapReduce as examples. However, there are many other open source options for Hadoop filesystems and data processing.
HDFS is a filesystem used to prevent data loss similar to the way RAID prevents it; however, HDFS is a bit different. HDFS by default stores three copies of each data block in the cluster on different nodes in the ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access