book

Text Mining with R

by Julia Silge, David Robinson

June 2017

Intermediate to advanced

191 pages

4h 29m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
OutlineTopics This Book Does Not CoverAbout This BookConventions Used in This BookUsing Code ExamplesO’Reilly SafariHow to Contact UsAcknowledgements
1. The Tidy Text Format
Contrasting Tidy Text with Other Data StructuresThe unnest_tokens FunctionTidying the Works of Jane AustenThe gutenbergr PackageWord FrequenciesSummary
2. Sentiment Analysis with Tidy Data
The sentiments DatasetSentiment Analysis with Inner JoinComparing the Three Sentiment DictionariesMost Common Positive and Negative WordsWordcloudsLooking at Units Beyond Just WordsSummary
3. Analyzing Word and Document Frequency: tf-idf
Term Frequency in Jane Austen’s NovelsZipf’s LawThe bind_tf_idf FunctionA Corpus of Physics TextsSummary
4. Relationships Between Words: N-grams and Correlations
Tokenizing by N-gramCounting and Filtering N-gramsAnalyzing BigramsUsing Bigrams to Provide Context in Sentiment AnalysisVisualizing a Network of Bigrams with ggraphVisualizing Bigrams in Other TextsCounting and Correlating Pairs of Words with the widyr PackageCounting and Correlating Among SectionsExamining Pairwise CorrelationSummary
5. Converting to and from Nontidy Formats
Tidying a Document-Term MatrixTidying DocumentTermMatrix ObjectsTidying dfm ObjectsCasting Tidy Text Data into a MatrixTidying Corpus Objects with MetadataExample: Mining Financial ArticlesSummary
6. Topic Modeling
Latent Dirichlet AllocationWord-Topic ProbabilitiesDocument-Topic ProbabilitiesExample: The Great Library HeistLDA on ChaptersPer-Document ClassificationBy-Word Assignments: augmentAlternative LDA ImplementationsSummary
7. Case Study: Comparing Twitter Archives
Getting the Data and Distribution of TweetsWord FrequenciesComparing Word UsageChanges in Word UseFavorites and RetweetsSummary
8. Case Study: Mining NASA Metadata
How Data Is Organized at NASAWrangling and Tidying the DataSome Initial Simple ExplorationWord Co-ocurrences and CorrelationsNetworks of Description and Title WordsNetworks of KeywordsCalculating tf-idf for the Description FieldsWhat Is tf-idf for the Description Field Words?Connecting Description Fields to KeywordsTopic ModelingCasting to a Document-Term MatrixReady for Topic ModelingInterpreting the Topic ModelConnecting Topic Modeling with KeywordsSummary
9. Case Study: Analyzing Usenet Text
PreprocessingPreprocessing TextWords in NewsgroupsFinding tf-idf Within NewsgroupsTopic ModelingSentiment AnalysisSentiment Analysis by WordSentiment Analysis by MessageN-gram AnalysisSummary

Bibliography
Index

Content preview from Text Mining with R

Chapter 4. Relationships Between Words: N-grams and Correlations

So far we’ve considered words as individual units, and considered their relationships to sentiments or to documents. However, many interesting text analyses are based on the relationships between words, whether examining which words tend to follow others immediately, or words that tend to co-occur within the same documents.

In this chapter, we’ll explore some of the methods tidytext offers for calculating and visualizing relationships between words in your text dataset. This includes the token = "ngrams" argument, which tokenizes by pairs of adjacent words rather than by individual ones. We’ll also introduce two new packages: ggraph, by Thomas Pedersen, which extends ggplot2 to construct network plots, and widyr, which calculates pairwise correlations and distances within a tidy data frame. Together these expand our toolbox for exploring text within the tidy data framework.

Tokenizing by N-gram

We’ve been using the unnest_tokens function to tokenize by word, or sometimes by sentence, which is useful for the kinds of sentiment and frequency analyses we’ve been doing so far. But we can also use the function to tokenize into consecutive sequences of words, called n-grams. By seeing how often word X is followed by word Y, we can then build a model of the relationships between them.

We do this by adding the token = "ngrams" option to unnest_tokens(), and setting n to the number of words we wish to capture in each n-gram. ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781491981641Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Text Mining with R

by Julia Silge, David Robinson

Chapter 4. Relationships Between Words: N-grams and Correlations

Tokenizing by N-gram

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.