Name
handle_starttag
Synopsis
h.handle_starttag(tag, attributes)Called to handle tags. tag is the tag
string, lowercased. attributes is a list
of pairs
(
name,value
),
where name is each
attribute’s name, lowercased, and
value is the value, processed to resolve
entity references and character references and to remove surrounding
quotes. HTMLParser’s
implementation of handle_starttag does nothing.
The following example uses HTMLParser to perform
the same task as our previous examples: fetching a page from the Web
with urllib, parsing it, and outputting the
hyperlinks.
import HTMLParser, urllib, urlparse
class LinksParser(HTMLParser.HTMLParser):
def __init_ _(self):
HTMLParser.HTMLParser.__init_ _(self)
self.seen = {}
def handle_starttag(self, tag, attributes):
if tag != 'a': return
for name, value in attributes:
if name == 'href' and value not in self.seen:
self.seen[value] = True
pieces = urlparse.urlparse(value)
if pieces[0] != 'http': return
print urlparse.urlunparse(pieces)
return
p = LinksParser( )
f = urllib.urlopen('http://www.python.org/index.html')
BUFSIZE = 8192
while True:
data = f.read(BUFSIZE)
if not data: break
p.feed(data)
p.close( )This example is similar to the one for sgmllib.
However, since the HTMLParser.HTMLParser
superclass performs no per-tag dispatching to methods, class
LinksParser needs to override method
handle_starttag and check that the
tag is indeed
'a‘.
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access