Skip to Content
View all events

Strata Data Superstream Series: Creating Data-Intensive Applications

Published by O'Reilly Media, Inc.

Intermediate content levelIntermediate
This live event utilizes interactive environmentsThis live event utilizes Jupyter Notebook technology

As the scale of data continues to grow (alongside an ever expanding ecosystem of tools to work with it), developing successful applications is an increasingly challenging proposition—and a necessity. At each stage of the process, from architecting to processing and storing data to deployment, there are a range of aspects to consider. Things like scalability, consistency, reliability, efficiency, and maintainability. It can be hard to figure out the right way forward.

In this event, you’ll gain insight into design and engineering best practices through interactive sessions and live coding demos. Join us to learn how to make the right decisions for your applications.

About the Strata Data Superstream Series: This four-part series of half-day online events gives attendees an overarching perspective of key topics that will help your organization maximize the business impact of your data.

What you’ll learn and how you can apply it

  • Learn how to build, scale, and test robust data pipelines
  • Get an overview of next-generation technologies for developing underlying data architectures
  • Explore design considerations to make your applications successful and reliable
  • Find out how AWS is empowering data professionals with a fully integrated set of tools
  • Learn how SAP approaches data preparation and orchestration at scale
  • Understand data governance techniques that you can implement right away

This live event is for you because...

  • You need to know the latest trends in data workflows, techniques, and tools.
  • You want to improve the scalability, reliability, security, and maintainability of your applications.
  • You want to better understand the systems you already use and the distributed systems on which modern databases are built.
  • You want to learn the best ways to integrate data governance into your applications.

Prerequisites

  • Come with your questions
  • Have a pen and paper handy to capture notes, insights, and inspiration

Recommended follow-up:

Schedule

The time frames are only estimates and may vary according to how the class is progressing.

Alistair Croll: Introduction (5 minutes) - 9:00am PT | 12:00pm ET | 4:00pm UTC/GMT

  • Alistair Croll welcomes you to the Strata Data Superstream.

Andrew McAfee and Alistair Croll: Fireside Chat—The Business and Societal Consequences of Data at Scale (30 minutes) - 9:05am PT | 12:05pm ET | 4:05pm UTC/GMT

  • Andrew McAfee joins Alistair Croll for a conversation on how data at scale changes business norms. They’ll touch on the significance of scarcity and abundance in the digital world, mistakes organizations make on the way to digital transformation, scaling up versus building for scale from the start, and more.
  • Andrew McAfee is the cofounder and codirector of MIT’s Initiative on the Digital Economy and a principal research scientist at the MIT Sloan School of Management. He studies how digital technologies are changing the world. His books include More from Less: The Surprising Story of How We Learned to Prosper Using Fewer Resources—and What Happens Next and the award-winning The Second Machine Age: Work, Progress, and Prosperity in a Time of Brilliant Technologies, coauthored with Erik Brynjolfsson. Andrew was educated at MIT and Harvard. He lives in Cambridge, MA, where he watches too much Red Sox baseball, doesn't ride his motorcycle enough, and starts his weekends with the New York Times’ Saturday crossword puzzle.

Sev Leonard: Prototype to Pipeline—Evolving from Data Exploration to Automated Data Processing (50 minutes) - 9:35am PT | 12:35pm ET | 4:35pm UTC/GMT

  • Great data-intensive applications start with ideas. Data pipelines come later, after you’ve gained an understanding of available data and how to transform and organize it. Join Sev Leonard to explore the considerations for separating data processing into pipeline steps, see how to identify opportunities for code reuse, and examine architectural considerations for scaling in size and breadth, accommodating both larger amounts of data and new data sources. Starting with just an idea, you’ll create a single source of information from API and text sources and learn how to move from data exploration to data pipelines.
  • Sev Leonard is a technologist, educator, and devoted cat dad with over two decades of experience in the tech industry. Sev is a senior software engineer at Fletch.ai and has worked as an analog engineer on Intel’s Core microprocessors; a consultant, writer, and teacher through his company, The Data Scout; and a software developer building data systems of all sizes that enabled groundbreaking advances in cancer research, improved access for millions of Medicaid and CHIP beneficiaries, and more. He’s active in the Python community as a mentor and presenter at PyCon, PyCascades, and local meetup groups.
  • Break (5 minutes)

Jose Kunnackal and Jesse Gebhardt: Delivering Seamless Embedded Analytics Within Applications (Sponsored by AWS) (30 minutes) - 10:30am PT | 1:30pm ET | 5:30pm UTC/GMT

  • As businesses grow increasingly data driven, data analytics and business intelligence are becoming pervasive. By moving information closer to the primary user experience (e.g., applications, websites, or internal enterprise portals), data practitioners are looking to deliver insights to the users where they are. AWS is seeing a rise in demand for embedding data visualizations and insights. However, embedding analytics into applications can often be expensive, complicated, and time-consuming. Jose Kunnackal and Jesse Gebhardt share AWS’s perspective on delivering a seamless embedded-analytics experience within applications.
  • Jose Kunnackal is senior manager of product management for Amazon QuickSight. Jose started his career at Motorola, writing software for telecom and first responder systems. Later he was director of engineering at Trilibis Mobile, where he built a SaaS mobile web platform using AWS services. Jose’s excited by the potential of cloud technologies and looks forward to helping customers move their analytics to the cloud.
  • Jesse Gebhardt is a principal specialist for Amazon QuickSight. He’s spent his entire career in business intelligence, at both Tableau and AWS. Jesse lives in sunny Phoenix, AZ. Outside of work, he’s an amateur electronic music producer and enjoys romping around with his two-year-old daughter.
  • This session will be followed by an hour-long Q&A in a breakout room. Stop by if you have more questions for Jose and Jesse.

Maureen Teyssier: Develop Data Science Products That Succeed (30 minutes) - 11:00am PT | 2:00pm ET | 6:00pm UTC/GMT

  • The majority of data science projects fail, especially products using machine learning and AI running in high-volume pipelines. To succeed, processes, like MLflow, have to live within a larger strategic context. Maureen Teyssier walks you through a logical construct that has succeeded again and again for a range of products and teams. If you’re a business or technical leader, you’ll learn how to identify key decisions that need to be made, who should be in the room for those decisions, and the appropriate development phases. And if you’re a developer, you’ll gain insight into what you should look for in a team and discover the additional contributions you can make to create stronger products.
  • Maureen Teyssier is chief data scientist at Reonomy, a commercial real estate data company transforming the world’s largest asset class with machine learning and AI. Maureen has run simulations and transformed data for 20 years. She has a breadth of knowledge on varieties of data, including location data, click data, image data, streaming data, and public and simulated data, as well as experience working with data at scale, managing datasets ranging from kilobytes to petabytes. She holds a PhD in computational astrophysics from Columbia University: she studied the evolution of galaxies by running cosmological simulations on supercomputers.
  • Break (5 minutes)

Kevin Poskitt and Silvio Arcangeli: The Evolution of Data Integration to Data Orchestration—Preparing and Integrating Data at Enterprise Scale (Sponsored by SAP) (30 minutes) - 11:35am PT | 2:35pm ET | 6:35pm UTC/GMT

  • Data-driven applications require advanced technologies with best practices to create agile, smarter business processes that are more resilient, profitable, and sustainable. Kevin Poskitt and Silvio Arcangeli explore the evolution of enterprise data integration to data orchestration and show how the requirements of data integration have stretched beyond the traditional approaches that worked so well in the past. If you’re struggling to catalog, integrate, prepare, and process your data at enterprise scale, this session is for you.
  • Kevin Poskitt is coauthor of the O’Reilly report Managing Data Orchestration & Integration at Scale: The Role of Open Source in Transforming Data. He’s senior director for SAP Data Intelligence, focused on product, go-to-market, and market strategy. He has a long history of working with various SAP technologies, from analytics to artificial intelligence to data integration and orchestration. Kevin has a passion for technology and data.
  • Silvio Arcangeli is a senior director for SAP Data Intelligence. An international technologist and enterprise architect with a deep and extensive technical background, he’s an advocate for business and people development and is passionate about enabling customers, evangelizing products, and presenting at public events.
  • This session will be followed by a 30-minute Q&A in a breakout room. Stop by if you have more questions for Kevin and Silvio.

Vinoo Ganesh: Watch Me Learn—Querying Data the Right Way (20 minutes) - 12:05pm PT | 3:05pm ET | 7:05pm UTC/GMT

  • The world has been flooded with data. Datasets that were megabytes have grown to terabytes, and datasets that were local and in-memory have become distributed across hundreds of boxes. But while the data itself has grown, expectations around it haven't changed. Data scientists, analysts, architects, and engineers still need to build fast applications to house, process, and analyze the data in a performant and effective way. And building data applications now requires an understanding of both the data and how it will be queried. Join Vinoo Ganesh to learn how cutting-edge data technologies work to optimize query performance. Using interactive scenarios, you’ll explore the optimization mechanisms that these applications use and learn how to leverage their features to write faster queries and data applications.
  • Vinoo Ganesh is chief technology officer at Veraset, a data-as-a-service startup focused on understanding the world from a geospatial perspective. Veraset processes, cleanses, and delivers over 3 TB of data daily and maintains incredibly strict uptime requirements. Vinoo previously managed the compute team at Palantir Technologies, tasked with managing Spark and its interaction with HDFS, S3, Parquet, YARN, and Kubernetes across the company.
  • Break (5 minutes)

Michele Goetz: Trends, Transitions, and Technical Advances—The 3 Ts of Data to Speed Up Apps in the Intelligent Edge (50 minutes) - 12:30pm PT | 3:30pm ET | 7:30pm UTC/GMT

  • Data is not exhaust, oil, or an asset. Data is the solution! In a world where chaos is the only constant, data speeds up apps and kicks the flywheel on business outcomes and resilience. Michele Goetz leads a dive into the three Ts—trends, transitions, and technical advances—emerging around edge intelligence, data mesh, and data marketplaces, showing you how you exploit them for the next killer app.
  • Michele Goetz is a vice president and principal analyst at Forrester Decisions, where she helps leaders navigate the complexities of data and differentiate with edge intelligence. Her research covers strategy, practice, and architecture for artificial intelligence, information management, data governance, and data ops.

Jessi Ashdown: Designing for Data Governance—Three Strategies You Can Implement Today (30 minutes) - 1:20pm PT | 4:20pm ET | 8:20pm UTC/GMT

  • You already know you need better data governance to meet regulations and make better data-driven business decisions. But it’s not easy. Existing frameworks for implementing governance may require more resources or a bigger budget than you have. Jessi Ashdown shares three basic strategies—minimal governance, ownership, and ongoing training—that you can implement to improve your current governance program and can incorporate into application design right away, with minimal cost and time.
  • Jessi Ashdown is a user experience researcher specializing in data governance at Google Cloud and a coauthor of O'Reilly's Data Governance: The Definitive Guide. She combines her background in research and passion for user experience to shed new light on interesting and innovative ways to successfully approach data governance in a modern world.

Alistair Croll: Closing Remarks (5 minutes) - 1:50pm PT | 4:50pm ET | 8:50pm UTC/GMT

  • Alistair Croll closes out today’s event.

Upcoming Strata Data Superstream events:

  • Data Warehouses, Data Lakes, and Data Lakehouses - August 10, 2021
  • Business Analysis - November 9, 2021

Your Host

  • Alistair Croll

    Alistair Croll is an entrepreneur, author, and conference organizer. He's written four books on technology and society, including the best-selling Lean Analytics, which has been translated into eight languages. He's the cofounder of web performance startup Coradiant (acquired by BMC), the Year One Labs startup accelerator, and a number of other early-stage companies.

    A prolific speaker, Alistair was a visiting executive at Harvard Business School, where he helped create a course on data science and critical thinking. He's founded and chaired a number of the world's leading technology events including Cloud Connect, Strata, Startupfest, Scaletech, and the FWD50 Digital Government conference. He's currently working on Just Evil Enough, the subversive marketing playbook. Alistair lives in Montreal, Canada, and writes at acroll.substack.com.

Skill covered

Data

Sponsored by

  • AWS logo
  • SAP logo