Scala for Data ScientistsThe Spark Programming ModelRecord LinkageGetting Started: The Spark Shell and SparkContextBringing Data from the Cluster to the ClientShipping Code from the Client to the ClusterStructuring Data with Tuples and Case ClassesAggregationsCreating HistogramsSummary Statistics for Continuous VariablesCreating Reusable Code for Computing Summary StatisticsSimple Variable Selection and ScoringWhere to Go from Here