June 2017
Beginner to intermediate
296 pages
7h 4m
English
One thing I want to point out at this point is that we're doing some pretty heavy lifting here. That self-join is a huge operation that generates a huge RDD. It's time to start splitting this up across more than just one CPU. Odds are your PC at home has at least two cores on it. Do this following little trick here; instead of just saying setMaster("local"), do setMaster("local[*]"). This syntax means, go take advantage of every core you have available on this computer and run a separate executor:
conf = SparkConf().setMaster("local[*]").setAppName("MovieSimilarities")
sc = SparkContext(conf = conf)
So for the first time, we're going to be running a Spark program across multiple executors. It will still be just on your ...
Read now
Unlock full access