Data Superstream: Data Engineering in the Age of AI
Published by O'Reilly Media, Inc.
Tools, Techniques, and Best Practices for Supporting AI and ML Use Cases
With the popularity of GenAI increasing by the day, good data engineering is more essential than ever. AI is transforming industries, but without high-quality data, even the most sophisticated models and AI agents can go awry. Data engineering is the backbone of AI workloads, as data engineers maintain and implement the infrastructure that houses data required for model training, build data pipelines that ensure accurate and accessible data, and monitor and enable model deployment.
Connect with leading-edge data engineering specialists to learn how to support AI and ML use cases, understand the emerging technologies you'll need to bridge the gap between data engineering and AI development, and pick up the latest best practices for dealing with common challenges and pitfalls.
What you’ll learn and how you can apply it
- Get an overview on how AI is bringing about a change in data engineering
- Learn how to use knowledge infrastructures and the Ontology Pipeline to build AI agents and systems
- Explore practical approaches for building streaming infrastructure and integrating AI agents using MCP
- Find out how to build scalable Gen AI pipelines with Spark NLP
- Discover the importance of open table formats, catalogs, and embedded systems as means for effective, governed Al development
- Understand the ethical considerations of working with data and AI
Recommended follow-up:
- Read Fundamentals of Data Engineering (book)
- Read AI Engineering (book)
- Read AI Agents with MCP (book)
- Take Unlock AI Superpowers for Data Pros (live course with Thomas Nield)
- Take Solving Data Preparation Tasks with AI (on-demand course)
- Take Generative AI for Data Management and Remediation (on-demand course)
Schedule
The time frames are only estimates and may vary according to how the class is progressing.
Introduction – Matt Housley (5 minutes)
- Matt welcomes you to the Data Superstream.
Keynote: Decentralized Data for a Centralized Brain – Valliappa Lakshmanan (30 minutes)
- Agentic AI is flipping the big data paradigm, demanding we bring the data to specialized intelligent compute, not the other way around. This shift fundamentally alters our assumptions about data modeling and storage, as LLMs perform in-context learning with vastly smaller datasets than traditional machine learning. Consequently, the growing context windows and tool calling capabilities in modern AI are quickly rendering many traditional ETL/ELT pipelines obsolete, forcing data engineers to radically rethink their entire approach. Valliappa Lakshmanan, cofounder and CTO of Obin AI, takes you through this paradigm shift, explores its future trajectory, and discusses how data engineers can adapt and thrive in this changing environment.
Designing Data Infrastructure in the Age of Generative AI – Lisa N. Cao (30 minutes)
- Developing powerful Al tooling has been the theme of the year, with agents and foundational models picking up steam across the board. But how do we serve data for agents to work effectively? What sort of interfaces and service mesh infrastructures will be required? What about at enterprise scale? What is context? Lisa N. Cao, developer relations and open source expert at Databricks, discusses the current big data landscape, the challenges to data platforming for Al, and the shifting importance of open table formats, catalogs, and embedded systems for effective, governed Al development. She’ll use open source technologies such as Apache Spark™, Unity Catalog OSS, and Apache Iceberg™ as key components of such reference architecture.
Break (5 minutes)
Knowledge Infrastructures and the Ontology Pipeline for AI Systems – Jessica Talisman (30 minutes)
- How do we build reliable, accurate AI agents and systems that scale across AI workloads? How can we engineer high-quality data pipelines delivering context-rich knowledge ecosystems for both humans and machines? Jessica Talisman, semantic architecture expert, provides some answers. Her Ontology Pipeline is a knowledge graph-based framework for designing and constructing semantic knowledge infrastructures powering AI. Its iterative logical steps—from controlled vocabularies, taxonomies, and metadata schemas to thesauri, ontologies, and knowledge graphs—harmonize data from diverse sources into structured, AI-ready assets. Discover how ontologies resolve semantic heterogeneity, boost interoperability, accelerate advanced AI analytics, and drive knowledge discovery. You’ll also explore real-world use cases and best practices for transforming unstructured and syntactic data into semantically rich, actionable assets underpinning AI workloads.
Building Scalable GenAI Inference Pipelines with Spark NLP – David Talby (30 minutes)
- David Talby, CEO of John Snow Labs and Pacific AI, looks at recent advances in the open source Spark NLP library that enable data engineers to build scalable batch GenAI inference pipelines. The emphasis is on scenarios where Spark provides a natural fit, such as processing multimodal corpora (documents, images, text) and precomputing embeddings or vision-language outputs for very large datasets. You’ll look at how Spark NLP enables efficient hardware utilization across CPUs, GPUs, and Apple Silicon, and how acceleration and quantization affect throughput in practice. You’ll also hear about integration with open-weight LLMs via Llama.cpp for private deployments, and the operational advantages of keeping ETL, enrichment, and inference within a single Spark-native workflow.
Break (5 minutes)
From Data Pipelines to AI Agents: Engineering the Future of Intelligent Systems – Scott Haines (30 minutes)
- Data engineering is evolving beyond traditional ETL pipelines. While speed and freshness remain critical, the real transformation lies in how we architect systems that don't just move data but rather enable AI agents to act on it intelligently in real-time. Software engineer Scott Haines, who specializes in distributed data systems and streaming technologies, explores practical approaches to building resilient streaming infrastructure with reduced operational overhead, and demonstrates how to integrate AI agents using Model Context Protocol (MCP) to create persistent memory streams that learn and adapt from your data flows.
The Evolving Data Engineer: Shaping AI, Shaped by AI – Lena Hall (30 minutes)
- A new era of intelligence is here. But who’ll actually build, manage, and ensure these advanced AI systems work correctly and reliably? The traditional lines of engineering are blurring, demanding skill sets that are both technically deeper and broader in scope than ever before. The engineer of tomorrow will be orchestrating complex ecosystems of specialized AI agents, rigorously validating their outputs and troubleshooting the toughest problems. Lena Hall, a senior director of developer relations at Akamai Technologies, ushers you into the future of data engineering—now inextricably entwined with AI. Join Lena to understand the evolving challenges and opportunities and how you can become a key player in building the future of intelligent systems.
Break (5 minutes)
Engineering Ethics in the Age of AI: Applying Ethical Frameworks to Data Practice – Mark Theunissen (30 minutes)
- Everyone is aware of the various moral risks that come with innovation and new technology. The emergence of big data and AI has supercharged these concerns—bringing questions of ethics, responsibility, and justice to the forefront. Educator and researcher Mark Theunissen introduces two leading ethical frameworks—responsible research and innovation (RRI) and value-sensitive design (VSD)—and considers their relevance for data engineers today. You’ll explore practical applications of these frameworks and gain insights into how engineers can navigate ethical complexity in their day-to-day work.
Closing Remarks – Matt Housley (5 minutes)
- Matt Housley closes out today’s event.
Your Hosts and Selected Speakers
Matt Housley
Matt Housley, a data engineering consultant and cloud specialist, is cofounder of Ternary Data, where he leverages his teaching experience to train future data engineers and advise teams on robust data architecture. After some early programming experience with Logo, Basic, and 6502 assembly, he completed a PhD in mathematics at the University of Utah. Matt then began working in data science, eventually specializing in cloud-based data engineering. Matt, Joe Reis, and their guests pontificate on all things data on The Monday Morning Data Chat.
Lisa N. Cao
Lisa N. Cao is an engineering, product, and advocacy expert in open source data infrastructure and DataOps fields. At Databricks, she oversees the open source involvement and developer relations of projects including MLflow, Apache Spark™, Delta Lake, Apache Iceberg™, and Unity Catalog OSS. Lisa also serves on the LF AI & Data governing board, formerly led the Open Platform for Enterprise AI’s (OPEA) developer experience working group, and leads the Continuous Delivery Foundation’s (CDF) DataOps Initiative.
Jessica Talisman
Jessica Talisman is a semantic architecture expert with over 25 years of experience designing information and knowledge frameworks for enterprise technology and cultural institutions. Drawing on her library science and learning sciences background, she creates high-integrity data ecosystems that blend information architecture, semantic systems, design thinking, and ethics frameworks. Jessica has held senior roles at Amazon, Adobe, Pluralsight, and GDIT, and founded Ontology Pipeline to offer consulting, courses, and collaborative media advancing structured-knowledge practices.
David Talby
David Talby is CEO of John Snow Labs and Pacific AI, helping companies apply artificial intelligence to solve real-world problems in healthcare and life science. He has extensive experience building and running web-scale software platforms and teams in startups, open source projects, and at Microsoft and Amazon. David holds a PhD in computer science and master’s degrees in computer science and business administration. He was named US CTO of the Year by the Global 100 Awards in 2022, Game Changers Awards in 2023, and ACQ5 Global Awards in 2025.
Scott Haines
Scott Haines is a seasoned software engineer specializing in massive distributed data systems and streaming technologies. Over the past decade, he has built and scaled data infrastructure at leading companies including Yahoo!, Twilio, Nike, and Buf. Scott is the author of books on Apache Spark and Delta Lake, and helps organizations successfully adopt open source technologies through his teaching and consulting work.
Valliappa Lakshmanan
Dr. Valliappa (“Lak”) Lakshmanan is cofounder and CTO of Obin AI, a startup that’s building deep domain AI agents for finance. He sets the technology and science direction of the company and is responsible for building the product. Lak is coauthor of several O’Reilly books, including Generative AI Design Patterns and Data Governance: The Definitive Guide.
Mark Theunissen
Mark Theunissen has a PhD in philosophy from the New School for Social Research. His work ranges from the philosophy of data, AI ethics, and philosophy of technology to philosophy of design and museum studies. Mark has worked as lecturer at NYU, Parsons School of Design, and Delft University of Technology, teaching courses on design thinking, innovation, philosophy of science, and engineering ethics. He has also worked as a curatorial researcher for the Jones Beach Energy & Nature Center.
Lena Hall
Lena Hall is a senior director of developer relations at Akamai Technologies with over 15 years of experience as a solution architect and technical leader. An expert in practical AI adoption, data engineering, cloud, and pragmatic architecture, she has led large, high-performing technical teams at AWS, Microsoft, and other companies. She’s also a cofounder of Droid AI, which helps people get real results from AI. Lena shares practical knowledge on her YouTube channel (Youtube.com/@lena-hall), LinkedIn (Linkedin.com/in/lena-hall), and at industry conferences as an international keynote speaker.