
4.2
widyr
パッケージによる
2
つの単語の出現頻度と相関
■
73
4.2.1
節単位の出現頻度と相関
第
2
章のセ ン チメント 分 析のとき と 同じよう に、『高慢 と 偏見』を
10
行の節
(
section
)に分割してみましょう。そのあとで、同じ節に共起する単語を調べます。
austen_section_words <- austen_books() %>%
filter(book == "Pride & Prejudice") %>%
mutate(section = row_number() %/% 10) %>%
filter(section > 0) %>%
unnest_tokens(word, text) %>%
filter(!word %in% stop_words$word)
austen_section_words
## # A tibble: 37,240
×
3
##
book section word
## <fctr> <dbl> <chr>
## 1 Pride & Prejudice 1 truth
## 2 Pride & Prejudice 1 universally
## 3 Pride & Prejudice 1 acknowledged
## 4 Pride & Prejudice 1 single
## 5 Pride & Prejudice 1 possession
## 6 Pride & Prejudice 1 fortune
## 7 Pride ...