November 2018
Beginner to intermediate
182 pages
4h 48m
English
Text and language is inherently unstructured. We might want to clean it in certain ways, such as expanding abbreviations and acronyms, removing punctuation, and so on. We also want to select a few samples that are the best representatives of the data we might see in the wild.
The other common practice is to prepare a gold dataset. A gold dataset is the best available data under reasonable conditions. This is not the best available data under ideal conditions. Creating the gold dataset often involves manual tagging and cleaning processes.
The next few sections are dedicated to text cleaning and text representations at this stage of the NLP workflow.
Read now
Unlock full access