Text Mining
Abstract
This chapter provides a detailed look into the emerging area of text mining and text analytics. It starts with a background of the origins of text mining and provides the motivation for this fascinating topic using the example of IBM's Watson, the Jeopardy—the winning computer program that was built almost entirely using concepts from text and data science. This chapter introduces some key concepts important in the area of text analytics such as term frequency–inverse document frequency scores. Finally, it describes two hands-on case studies in which it is shown how to use RapidMiner to address problems like document clustering and automatic gender classification based on text content.
Keywords
Inverse document frequency; ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access