Foreword by Alex Ratner
The real-world impact of artificial intelligence (AI) has grown substantially in recent years, largely due to the advent of deep learning models. These models are more powerful and push-button than ever before—learning their own powerful, distributed representations directly from raw data with minimal to no manual feature engineering across diverse data and task types. They are also increasingly commoditized and accessible in the open source.
However, deep learning models are also more data hungry than ever, requiring massive, carefully labeled training datasets to function. In a world where the latest and greatest model architectures are downloadable in seconds and the powerful hardware needed to train them is a click away in the cloud, access to high-quality labeled training data has become a major differentiator across both industry and academia. More succinctly, we have left the age of model-centric AI and are entering the era of data-centric AI.
Unfortunately, labeling data at the scale and quality required to train—or “supervise”—useful AI models tends to be both expensive and time-consuming because it requires manual human input over huge numbers of examples. Person-years of data labeling per model is not uncommon, and when model requirements change—say, to classify medical images as “normal,” “abnormal,” or “emergent” rather than just “normal” or “abnormal”—data must often be relabeled from scratch. When organizations are deploying tens, hundreds, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access