Start with simple rules
I have written a script which scrapes the Wikipedia page entitled Programming language.
Click here to open that page: https://en.wikipedia.org/wiki/Programming_language
Extracting the name of the programming languages from the text of the given page is our goal. Take an example: The page has C, C++, Java, JavaScript, and so on, programming languages. I want to extract them. These words can be a part of sentences or have occurred standalone in the text data content.
Now, see how we can solve this problem by defining a simple rule. The GitHub link for the script is: https://github.com/jalajthanaki/NLPython/blob/master/ch7/7_1_simplerule.py
The data file link on GitHub is: https://github.com/jalajthanaki/NLPython/blob/master/data/simpleruledata.txt ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access