book

Text Mining with R

by Julia Silge, David Robinson

June 2017

Intermediate to advanced

191 pages

4h 29m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
OutlineTopics This Book Does Not CoverAbout This BookConventions Used in This BookUsing Code ExamplesO’Reilly SafariHow to Contact UsAcknowledgements
1. The Tidy Text Format
Contrasting Tidy Text with Other Data StructuresThe unnest_tokens FunctionTidying the Works of Jane AustenThe gutenbergr PackageWord FrequenciesSummary
2. Sentiment Analysis with Tidy Data
The sentiments DatasetSentiment Analysis with Inner JoinComparing the Three Sentiment DictionariesMost Common Positive and Negative WordsWordcloudsLooking at Units Beyond Just WordsSummary
3. Analyzing Word and Document Frequency: tf-idf
Term Frequency in Jane Austen’s NovelsZipf’s LawThe bind_tf_idf FunctionA Corpus of Physics TextsSummary
4. Relationships Between Words: N-grams and Correlations
Tokenizing by N-gramCounting and Filtering N-gramsAnalyzing BigramsUsing Bigrams to Provide Context in Sentiment AnalysisVisualizing a Network of Bigrams with ggraphVisualizing Bigrams in Other TextsCounting and Correlating Pairs of Words with the widyr PackageCounting and Correlating Among SectionsExamining Pairwise CorrelationSummary
5. Converting to and from Nontidy Formats
Tidying a Document-Term MatrixTidying DocumentTermMatrix ObjectsTidying dfm ObjectsCasting Tidy Text Data into a MatrixTidying Corpus Objects with MetadataExample: Mining Financial ArticlesSummary
6. Topic Modeling
Latent Dirichlet AllocationWord-Topic ProbabilitiesDocument-Topic ProbabilitiesExample: The Great Library HeistLDA on ChaptersPer-Document ClassificationBy-Word Assignments: augmentAlternative LDA ImplementationsSummary
7. Case Study: Comparing Twitter Archives
Getting the Data and Distribution of TweetsWord FrequenciesComparing Word UsageChanges in Word UseFavorites and RetweetsSummary
8. Case Study: Mining NASA Metadata
How Data Is Organized at NASAWrangling and Tidying the DataSome Initial Simple ExplorationWord Co-ocurrences and CorrelationsNetworks of Description and Title WordsNetworks of KeywordsCalculating tf-idf for the Description FieldsWhat Is tf-idf for the Description Field Words?Connecting Description Fields to KeywordsTopic ModelingCasting to a Document-Term MatrixReady for Topic ModelingInterpreting the Topic ModelConnecting Topic Modeling with KeywordsSummary
9. Case Study: Analyzing Usenet Text
PreprocessingPreprocessing TextWords in NewsgroupsFinding tf-idf Within NewsgroupsTopic ModelingSentiment AnalysisSentiment Analysis by WordSentiment Analysis by MessageN-gram AnalysisSummary

Bibliography
Index

Content preview from Text Mining with R

Chapter 7. Case Study: Comparing Twitter Archives

One type of text that gets plenty of attention is text shared online via Twitter. In fact, several of the sentiment lexicons used in this book (and commonly used in general) were designed for use with and validated on tweets. Both authors of this book are on Twitter and are fairly regular users of it, so in this case study, let’s compare the entire Twitter archives of Julia and David.

Getting the Data and Distribution of Tweets

An individual can download his or her own Twitter archive by following directions available on Twitter’s website. We each downloaded ours and will now open them up. Let’s use the lubridate package to convert the string timestamps to date-time objects and initially take a look at our tweeting patterns overall (Figure 7-1).

library(lubridate)
library(ggplot2)
library(dplyr)
library(readr)

tweets_julia <- read_csv("data/tweets_julia.csv")
tweets_dave <- read_csv("data/tweets_dave.csv")
tweets <- bind_rows(tweets_julia %>%
                      mutate(person = "Julia"),
                    tweets_dave %>%
                      mutate(person = "David")) %>%
  mutate(timestamp = ymd_hms(timestamp))

ggplot(tweets, aes(x = timestamp, fill = person)) +
  geom_histogram(position = "identity", bins = 20, show.legend = FALSE) +
  facet_wrap(~person, ncol = 1)

David and Julia tweet at about the same rate currently and joined Twitter about a year ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781491981641Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Text Mining with R

by Julia Silge, David Robinson

Chapter 7. Case Study: Comparing Twitter Archives

Getting the Data and Distribution of Tweets

Figure 7-1. All tweets from our accounts

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.