Skip to Content
View all events

Data Cleaning for AI: From Raw Data to Reliable Intelligence

Published by O'Reilly Media, Inc.

Beginner to intermediate content levelBeginner to intermediate

From raw data to reliable intelligence

What you’ll learn and how you can apply it

  • Profile any dataset to identify quality issues that would degrade ML model performance
  • Calculate data quality scores across six dimensions and set model-training thresholds
  • Apply deduplication, null-handling, and schema enforcement strategies specific to training data
  • Prepare datasets for feature stores and implement basic versioning for reproducible ML pipelines

Course description

The quality of your AI system is only as good as the data it learns from. Drawing from their textbook Data Cleaning and Data Quality: A Comprehensive Guide, Dr. James J. Foster and Emmanuel Jackson give data engineers, data scientists, analysts, and AI practitioners a deeper understanding of how to transform raw, production-grade data into the clean, validated, AI-ready datasets that modern machine learning systems demand.

You’ll discover the real-world data quality failures that silently degrade model accuracy, amplify bias, and erode trust in AI outputs. Across four modules, you’ll examine the specific problems that arise when incomplete customer records, inconsistent product catalogs, duplicate entities, and unstandardized fields enter a model pipeline, and you’ll learn the frameworks, scoring methods, and architectural patterns you need to prevent them.

This live event is for you because...

  • You’re a mid-to-senior-level data professional, analytics leader, or AI program manager who’s responsible for the quality and reliability of the data that powers machine learning and AI initiatives.
  • You’re an AI or ML engineer who owns the model development lifecycle and wants to better understand data quality dimensions, scoring methods, and governance structures.

Prerequisites

  • A basic understanding of data concepts
  • Awareness of data pipelines and ETL processes
  • General familiarity with AI and machine learning concepts
  • Experience working in or interacting with data teams (helpful but not required)

Recommended follow-up:

Schedule

The time frames are only estimates and may vary according to how the class is progressing.

Why AI amplifies the data quality crisis (45 minutes)

  • Presentation: Welcome to the AI data quality crisis—the truth about biased models, silent failures, and the $12.9M cost of bad data; understand the enemy—the six dimensions of data quality and hands-on profiling of sample_customers.Csv for AI readiness
  • Group discussion: What went wrong with the “train on everything” approach?
  • Q&A

How to profile, measure, and score your data (45 minutes)

  • Presentation: You can’t fix what you can’t measure—a five-step profiling process to prevent AI failures; how to use a balanced scorecard for AI data readiness; from scores to signals—quantifying each dimension, using green/amber/red thresholds as quality gates, and matching monitoring frequency to AI pipeline criticality
  • Group discussion: The challenge of measurement—getting buy-in for data quality KPIs
  • Q&A
  • Break

A structured approach to fixing bad data (45 minutes)

  • Presentation: The correct order of operations—why you must standardize and validate before you enrich and deduplicate; case study (87.5% reduction in transaction failures from 3.2% to 0.4%)
  • Group discussion: How to handle data that fails validation (quarantine or reject?)
  • Q&A

Governance, compliance, and the future (45 minutes)

  • Presentation: From manual fixes to enterprise capability—embedding quality in ETL pipelines, automating checks, establishing clear roles and accountability; how AI is transforming data cleaning, the demand for data quality professionals, and seven principles for AI-ready data
  • Q&A

Your Instructors

  • James C. Foster

    Dr. James J. Foster is a principal data scientist at Evanston Technology Partners with over a decade of experience applying machine learning and predictive analytics to complex enterprise data problems. He has built data quality and analytics programs across manufacturing, agriculture, and technology sectors. Foster has spoken at Predictive Analytics World Manufacturing, INFORMS, AIChE, and other industry conferences, and he codeveloped the Diamond Mines to Diamond Minds program, a large-scale data practitioner training initiative operating across two continents. His work sits at the intersection of data engineering rigor and applied ML, making him well-positioned to teach the specific data preparation skills AI practitioners need.

  • Emmanuel Jackson

    Emmanuel Jackson is CEO of Evanston Technology Partners, where he leads data engineering and technology training programs for enterprise clients. He cocreated the Diamond Mines to Diamond Minds program with James Foster, bringing the organizational and instructional design perspective to the partnership and a track record of translating complex data concepts into accessible, job-ready training. He has overseen the delivery of data skills training to practitioners across the United States and Africa.

Skills covered

  • Data Wrangling, Preparation, Cleaning
  • Machine Learning
  • Data Models
  • Data Quality