Skip to Content
Essential PySpark for Scalable Data Analytics
book

Essential PySpark for Scalable Data Analytics

by Sreeram Nudurupati
October 2021
Beginner to intermediate
322 pages
7h 27m
English
Packt Publishing
Content preview from Essential PySpark for Scalable Data Analytics

Chapter 4: Real-Time Data Analytics

In the modern big data world, data is being generated at a tremendous pace, that is, faster than any of the past decade's technologies can handle, such as batch processing ETL tools, data warehouses, or business analytics systems. It is essential to process data and draw insights in real time for businesses to make tactical decisions that help them to stay competitive. Therefore, there is a need for real-time analytics systems that can process data in real or near real-time and help end users get to the latest data as quickly as possible.

In this chapter, you will explore the architecture and components of a real-time big data analytics processing system, including message queues as data sources, Delta as ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Start your free trial

You might also like

Data Analytics with Hadoop

Data Analytics with Hadoop

Benjamin Bengfort, Jenny Kim
Data Science on AWS

Data Science on AWS

Chris Fregly, Antje Barth

Publisher Resources

ISBN: 9781800568877Supplemental Content