March 2019
Beginner
490 pages
12h 40m
English
In this script, we can see how to extract links using urllib and HTMLParser. HTMLParser is a module that allows us to parse text files formatted in HTML. You can get more information at https://docs.python.org/3/library/html.parser.html.
You can find the following code in the extract_links_parser.py file:
#!/usr/bin/env python3from html.parser import HTMLParserimport urllib.requestclass myParser(HTMLParser): def handle_starttag(self, tag, attrs): if (tag == "a"): for a in attrs: if (a[0] == 'href'): link = a[1] if (link.find('http') >= 0): print(link) newParse = myParser() newParse.feed(link)url = "http://www.packtpub.com"request = urllib.request.urlopen(url)parser = myParser()parser.feed(request.read().decode('utf-8')) ...Read now
Unlock full access