book

Text Mining with R

by Julia Silge, David Robinson

June 2017

Intermediate to advanced

191 pages

4h 29m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
OutlineTopics This Book Does Not CoverAbout This BookConventions Used in This BookUsing Code ExamplesO’Reilly SafariHow to Contact UsAcknowledgements
1. The Tidy Text Format
Contrasting Tidy Text with Other Data StructuresThe unnest_tokens FunctionTidying the Works of Jane AustenThe gutenbergr PackageWord FrequenciesSummary
2. Sentiment Analysis with Tidy Data
The sentiments DatasetSentiment Analysis with Inner JoinComparing the Three Sentiment DictionariesMost Common Positive and Negative WordsWordcloudsLooking at Units Beyond Just WordsSummary
3. Analyzing Word and Document Frequency: tf-idf
Term Frequency in Jane Austen’s NovelsZipf’s LawThe bind_tf_idf FunctionA Corpus of Physics TextsSummary
4. Relationships Between Words: N-grams and Correlations
Tokenizing by N-gramCounting and Filtering N-gramsAnalyzing BigramsUsing Bigrams to Provide Context in Sentiment AnalysisVisualizing a Network of Bigrams with ggraphVisualizing Bigrams in Other TextsCounting and Correlating Pairs of Words with the widyr PackageCounting and Correlating Among SectionsExamining Pairwise CorrelationSummary
5. Converting to and from Nontidy Formats
Tidying a Document-Term MatrixTidying DocumentTermMatrix ObjectsTidying dfm ObjectsCasting Tidy Text Data into a MatrixTidying Corpus Objects with MetadataExample: Mining Financial ArticlesSummary
6. Topic Modeling
Latent Dirichlet AllocationWord-Topic ProbabilitiesDocument-Topic ProbabilitiesExample: The Great Library HeistLDA on ChaptersPer-Document ClassificationBy-Word Assignments: augmentAlternative LDA ImplementationsSummary
7. Case Study: Comparing Twitter Archives
Getting the Data and Distribution of TweetsWord FrequenciesComparing Word UsageChanges in Word UseFavorites and RetweetsSummary
8. Case Study: Mining NASA Metadata
How Data Is Organized at NASAWrangling and Tidying the DataSome Initial Simple ExplorationWord Co-ocurrences and CorrelationsNetworks of Description and Title WordsNetworks of KeywordsCalculating tf-idf for the Description FieldsWhat Is tf-idf for the Description Field Words?Connecting Description Fields to KeywordsTopic ModelingCasting to a Document-Term MatrixReady for Topic ModelingInterpreting the Topic ModelConnecting Topic Modeling with KeywordsSummary
9. Case Study: Analyzing Usenet Text
PreprocessingPreprocessing TextWords in NewsgroupsFinding tf-idf Within NewsgroupsTopic ModelingSentiment AnalysisSentiment Analysis by WordSentiment Analysis by MessageN-gram AnalysisSummary

Bibliography
Index

Content preview from Text Mining with R

Chapter 3. Analyzing Word and Document Frequency: tf-idf

A central question in text mining and natural language processing is how to quantify what a document is about. Can we do this by looking at the words that make up the document? One measure of how important a word may be is its term frequency (tf), how frequently a word occurs in a document, as we examined in Chapter 1. There are words in a document, however, that occur many times but may not be important; in English, these are probably words like “the,” “is,” “of,” and so forth. We might take the approach of adding words like these to a list of stop words and removing them before analysis, but it is possible that some of these words might be more important in some documents than others. A list of stop words is not a very sophisticated approach to adjusting term frequency for commonly used words.

Another approach is to look at a term’s inverse document frequency (idf), which decreases the weight for commonly used words and increases the weight for words that are not used very much in a collection of documents. This can be combined with term frequency to calculate a term’s tf-idf (the two quantities multiplied together), the frequency of a term adjusted for how rarely it is used.

Note

The statistic tf-idf is intended to measure how important a word is to a document in a collection (or corpus) of documents, for example, to one novel in a collection of novels or to one website in a collection of websites.

The statistic tf-idf ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781491981641Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Text Mining with R

by Julia Silge, David Robinson

Chapter 3. Analyzing Word and Document Frequency: tf-idf

Note

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.