Summary
Spark Streaming is based on a micro-batching model that is suitable for applications with throughput and high latency (> 0.5 seconds). Spark Streaming's DStream API provides transformations and actions for working with DStreams, including conventional transformations, window operations, output actions, and stateful operations such as updateStateByKey. Spark Streaming supports a variety of input sources and output sources used in the Big Data ecosystem. Spark Streaming supports the direct approach with Kafka, which really provides great benefits such as exactly once processing and avoiding WAL replication.
There are two types of failures in a Spark Streaming application; executor failure and driver failure. Executor failures are automatically ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access