Word embeddings
Up until this point, we have used scikit-learn to embed documents (tweets, reviews, URLs, and so on) into a vectorized format by regarding tokens (words, n-grams) as features and documents as having a certain amount of these tokens. For example, if we had 1,583 documents and we told our CountVectorizer to learn the top 1,000 tokens of ngram_range from one to five, we would end up with a matrix of shape (1583, 1000) where each row represented a single document and the 1,000 columns represented literal n-grams found in the corpus. But how do we achieve an even lower level of understanding? How do we start to teach the machine what words mean in context?
For example, if we were to ask you the following questions, you may give ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access