Customizing parser tools
In real life, datasets are quite complex and messy. In that case, it may be that the parser is unable to give you a perfect or accurate result. Let's take an example.
Let's assume that you want to parse a dataset that has text content of research papers, and these research papers belong to the chemistry domain. If you are using the Stanford parser in order to generate a parse tree for this dataset, then sentences that contain chemical symbols and equations may not get parsed properly. This is because the Stanford parser has been trained on the Penn TreeBank corpus, so it's accuracy for the generation of a parse tree for chemical symbols and equation is low. In this case, you have two options - either you search for ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access