Negative sampling
Let's say we are building a CBOW model and we have a sentence Birds are flying in the sky. Let the context words be birds, are, in, and the and the target word be flying.
We need to update the weights of the network every time it predicts the incorrect target word. So, except for the word flying, if a different word is predicted as a target word, then we update the network.
But this is just a small set of vocabulary. Consider the case where we have millions of words in the vocabulary. In that case, we need to perform numerous weight updates until the network predict the correct target word. It is time-consuming and also not an efficient method. So, instead of doing this, we mark the correct target word as a positive class ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access