Skip to Main Content
Hadoop Security
book

Hadoop Security

by Ben Spivey, Joey Echeverria
June 2015
Intermediate to advanced content levelIntermediate to advanced
340 pages
8h 43m
English
O'Reilly Media, Inc.
Content preview from Hadoop Security

Chapter 1. Introduction

Back in 2003, Google published a paper describing a scale-out architecture for storing massive amounts of data across clusters of servers, which it called the Google File System (GFS). A year later, Google published another paper describing a programming model called MapReduce, which took advantage of GFS to process data in a parallel fashion, bringing the program to where the data resides. Around the same time, Doug Cutting and others were building an open source web crawler now called Apache Nutch. The Nutch developers realized that the MapReduce programming model and GFS were the perfect building blocks for a distributed web crawler, and they began implementing their own versions of both projects. These components would later split from Nutch and form the Apache Hadoop project. The ecosystem1 of projects built around Hadoop’s scale-out architecture brought about a different way of approaching problems by allowing the storage and processing of all data important to a business.

While all these new and exciting ways to process and store data in the Hadoop ecosystem have brought many use cases across different verticals to use this technology, it has become apparent that managing petabytes of data in a single centralized cluster can be dangerous. Hundreds if not thousands of servers linked together in a common application stack raises many questions about how to protect such a valuable asset. While other books focus on such things as writing MapReduce code, ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Start your free trial

You might also like

Introduction to Hadoop Security

Introduction to Hadoop Security

Jeff Bean
Securing Hadoop

Securing Hadoop

Sudheesh Narayan

Publisher Resources

ISBN: 9781491900970Errata Page