September 2016
Beginner to intermediate
326 pages
6h 47m
English
Datasets are similar to RDDs; however, instead of using Java or Kryo Serialization, they use a specialized Encoder to serialize the objects for processing or transmitting over the network. While both encoders and standard serialization are responsible for turning an object into bytes, encoders are generated dynamically and use a format that allows Spark to perform many operations such as filtering, sorting, and hashing without deserializing the bytes back into an object. Source: https://spark.apache.org/docs/latest/sql-programming-guide.html#creating-datasets.
The following Scala example creates a Dataset and DataFrame from an RDD. Enter the scala shell with the spark-shell command:
scala> case class ...Read now
Unlock full access