July 2017
Beginner to intermediate
486 pages
13h 49m
English
Subsampling is also one of the techniques that we use when we are building word pairs, and as we know, these word pairs are sample training data.
Subsampling is the method that removes the most frequent words. This technique is very useful for removing stop words.
These techniques also remove words randomly, and these randomly chosen words occur in the corpus more frequently. So, words that are removed are more frequent than some threshold t with a probability of p, where f marks the words corpus frequency and we use t = 10−5 in our experiments. Refer to the following equation given in Figure 6.15:
Read now
Unlock full access