book

Text Mining with R

by Julia Silge, David Robinson

June 2017

Intermediate to advanced

191 pages

4h 29m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
OutlineTopics This Book Does Not CoverAbout This BookConventions Used in This BookUsing Code ExamplesO’Reilly SafariHow to Contact UsAcknowledgements
1. The Tidy Text Format
Contrasting Tidy Text with Other Data StructuresThe unnest_tokens FunctionTidying the Works of Jane AustenThe gutenbergr PackageWord FrequenciesSummary
2. Sentiment Analysis with Tidy Data
The sentiments DatasetSentiment Analysis with Inner JoinComparing the Three Sentiment DictionariesMost Common Positive and Negative WordsWordcloudsLooking at Units Beyond Just WordsSummary
3. Analyzing Word and Document Frequency: tf-idf
Term Frequency in Jane Austen’s NovelsZipf’s LawThe bind_tf_idf FunctionA Corpus of Physics TextsSummary
4. Relationships Between Words: N-grams and Correlations
Tokenizing by N-gramCounting and Filtering N-gramsAnalyzing BigramsUsing Bigrams to Provide Context in Sentiment AnalysisVisualizing a Network of Bigrams with ggraphVisualizing Bigrams in Other TextsCounting and Correlating Pairs of Words with the widyr PackageCounting and Correlating Among SectionsExamining Pairwise CorrelationSummary
5. Converting to and from Nontidy Formats
Tidying a Document-Term MatrixTidying DocumentTermMatrix ObjectsTidying dfm ObjectsCasting Tidy Text Data into a MatrixTidying Corpus Objects with MetadataExample: Mining Financial ArticlesSummary
6. Topic Modeling
Latent Dirichlet AllocationWord-Topic ProbabilitiesDocument-Topic ProbabilitiesExample: The Great Library HeistLDA on ChaptersPer-Document ClassificationBy-Word Assignments: augmentAlternative LDA ImplementationsSummary
7. Case Study: Comparing Twitter Archives
Getting the Data and Distribution of TweetsWord FrequenciesComparing Word UsageChanges in Word UseFavorites and RetweetsSummary
8. Case Study: Mining NASA Metadata
How Data Is Organized at NASAWrangling and Tidying the DataSome Initial Simple ExplorationWord Co-ocurrences and CorrelationsNetworks of Description and Title WordsNetworks of KeywordsCalculating tf-idf for the Description FieldsWhat Is tf-idf for the Description Field Words?Connecting Description Fields to KeywordsTopic ModelingCasting to a Document-Term MatrixReady for Topic ModelingInterpreting the Topic ModelConnecting Topic Modeling with KeywordsSummary
9. Case Study: Analyzing Usenet Text
PreprocessingPreprocessing TextWords in NewsgroupsFinding tf-idf Within NewsgroupsTopic ModelingSentiment AnalysisSentiment Analysis by WordSentiment Analysis by MessageN-gram AnalysisSummary

Bibliography
Index

Overview

Much of the data available today is unstructured and text-heavy, making it challenging for analysts to apply their usual data wrangling and visualization tools. With this practical book, you’ll explore text-mining techniques with tidytext, a package that authors Julia Silge and David Robinson developed using the tidy principles behind R packages like ggraph and dplyr. You’ll learn how tidytext and other tidy tools in R can make text analysis easier and more effective.

The authors demonstrate how treating text as data frames enables you to manipulate, summarize, and visualize characteristics of text. You’ll also learn how to integrate natural language processing (NLP) into effective workflows. Practical code examples and data explorations will help you generate real insights from literature, news, and social media.

Learn how to apply the tidy text format to NLP
Use sentiment analysis to mine the emotional content of text
Identify a document’s most important terms with frequency measurements
Explore relationships and connections between words with the ggraph and widyr packages
Convert back and forth between R’s tidy and non-tidy text formats
Use topic modeling to classify document collections into natural groups
Examine case studies that compare Twitter archives, dig into NASA metadata, and analyze thousands of Usenet messages

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781491981641Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills