One-hot encoding
In an NLP application, you always get categorical data. The categorical data is mostly in the form of words. There are words that form the vocabulary. The words from this vocabulary cannot turn into vectors easily.
Consider that you have a vocabulary with the size N. The way to approximate the state of the language is by representing the words in the form of one-hot encoding. This technique is used to map the words to the vectors of length n, where the nth digit is an indicator of the presence of the particular word. If you are converting words to the one-hot encoding format, then you will see vectors such as 0000...001, 0000...100, 0000...010, and so on. Every word in the vocabulary is represented by one of the combinations ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access