May 2019
Intermediate to advanced
272 pages
7h 19m
English
We will focus on the 1-Billion Word dataset that was proposed in 2013 in the paper One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling by Ciprian Chelba et al. This dataset can be downloaded from http://www.statmt.org/lm-benchmark/.
We will start with the imports, which include the libraries and functions in this subsection:
from collections import defaultdict, Counterimport numpy as npfrom glob import globimport pandas as pd
Now, we are going to define the method that loads the text data from disk and preprocesses it. The function signature includes variables to handle the start of the sentence (sos), the end of the sentence (eos), and unknown (unk_token) tokens:
def get_data(glob_str, vocabulary_size=96, ...
Read now
Unlock full access