Hadoop for big data
Apache Hadoop is a 100 percent open source software framework used for two important fundamental tasks: storing and processing big data. It has been the leading big data tool for distributed parallel processing of data stored across multiple servers and is able to scale without limits. Because of its scalability, flexibility, fault tolerance, and low-cost features, many cloud-based solution vendors, financial institutions, and enterprises use Hadoop for their big data needs.
The Hadoop framework contains modules that are critical to its functions: the Hadoop Distributed File System (HDFS), Yet Another Resource Negotiator (YARN), and MapReduce (MapR).
HDFS
HDFS is a file system unique to Hadoop that is designed to be scalable and ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access