Skip to Content
Interactive Spark using PySpark
book

Interactive Spark using PySpark

by Benjamin Bengfort, Jenny Kim
August 2016
Intermediate to advanced
20 pages
41m
English
O'Reilly Media, Inc.

Overview

Apache Spark is an in-memory framework that allows data scientists to explore and interact with big data much more quickly than with Hadoop. Python users can work with Spark using an interactive shell called PySpark.

Why is it important?

PySpark makes the large-scale data processing capabilities of Apache Spark accessible to data scientists who are more familiar with Python than Scala or Java. This also allows for reuse of a wide variety of Python libraries for machine learning, data visualization, numerical analysis, etc.

What you'll learn—and how you can apply it

Compare the different components provided by Spark, and what use cases they fit. Learn how to use RDDs (resilient distributed datasets) with PySpark. Write Spark applications in Python and submit them to the cluster as Spark jobs. Get an introduction to the Spark computing framework. Apply this approach to a worked example to determine the most frequent airline delays in a specific month and year.

This lesson is for you because…

  • You're a data scientist, familiar with Python coding, who needs to get up and running with PySpark
  • You're a Python developer who needs to leverage the distributed computing resources available on a Hadoop cluster, without learning Java or Scala first

Prerequisites

  • Familiarity with writing Python applications
  • Some familiarity with bash command-line operations
  • Basic understanding of how to use simple functional programming constructs in Python, such as closures, lambdas, maps, etc.

Materials or downloads needed in advance

This lesson is taken from Data Analytics with Hadoop by Jenny Kim and Benjamin Bengfort.

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

Apache Spark with Python - Big Data with PySpark and Spark

Apache Spark with Python - Big Data with PySpark and Spark

James Lee

Publisher Resources

ISBN: 9781491965313Errata Page