O'Reilly logo

Machine Learning with Spark - Second Edition by Nick Pentreath, Manpreet Singh Ghotra, Rajdeep Dua

Stay ahead with the world's most comprehensive technology and business learning platform.

With Safari, you learn the way you learn best. Get unlimited access to videos, live online training, learning paths, books, tutorials, and more.

Start Free Trial

No credit card required

Advanced Text Processing with Spark

In Chapter 4, Obtaining, Processing, and Preparing Data with Spark, we covered various topics related to feature extraction and data processing, including the basics of extracting features from text data. In this chapter, we will introduce more advanced text processing techniques available in Spark ML to work with large-scale text datasets.

In this chapter, we will:

  • Work through detailed examples that illustrate data processing, feature extraction, and the modeling pipeline, as they relate to text data
  • Evaluate the similarity between two documents based on the words in the documents
  • Use the extracted text features as inputs for a classification model
  • Cover a recent development in natural language processing ...

With Safari, you learn the way you learn best. Get unlimited access to videos, live online training, learning paths, books, interactive tutorials, and more.

Start Free Trial

No credit card required