Introducing Apache Spark
If you have worked in big data, there is a high probability that you already know what Apache Spark is, and you can skip this section. But if you don't, don't worry—we'll go through the basics.
Spark is a powerful, fast, and scalable real-time data analytics engine for large scale data processing. It's an open source framework that was developed initially by the UC Berkeley AMPLab around the year 2009. Around 2013, AMPLab contributed Spark to the Apache Software Foundation, with Apache Spark Community releasing Spark 1.0 in 2014.
The community continues to make regular releases and brings new features into the project. At the time of writing this book, we have the Apache Spark 2.4.0 release and active community on ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access