Summary
In this chapter, we introduced a new class of algorithms, that is, frequent pattern mining applications, and showed you how to deploy them in a real-world scenario. We first discussed the very basics of pattern mining and the problems that can be addressed using these techniques. In particular, we saw how to implement the three available algorithms in Spark, FP-growth, association rules, and prefix span. As a running example for the applications we used clickstream data provided by MSNBC, which also helped us to compare the algorithms qualitatively.
Next, we introduced the basic terminology and entry points of Spark Streaming and considered a few real-world scenarios. We discussed how to deploy and evaluate one of the frequent pattern ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access