Counting Tags in a Document
Credit: Paul Prescod
Problem
You want to get a sense of how often particular elements occur in an XML document, and the relevant counts must be extracted rapidly.
Solution
You can subclass
SAX’s
ContentHandler to make your own specialized
classes for any kind of task, including the collection of such
statistics:
from xml.sax.handler import ContentHandler
import xml.sax
class countHandler(ContentHandler):
def _ _init_ _(self):
self.tags={}
def startElement(self, name, attr):
if not self.tags.has_key(name):
self.tags[name] = 0
self.tags[name] += 1
parser = xml.sax.make_parser( )
handler = countHandler( )
parser.setContentHandler(handler)
parser.parse("test.xml")
tags = handler.tags.keys( )
tags.sort( )
for tag in tags:
print tag, handler.tags[tag]Discussion
When I start with a new XML content set, I like to get a sense of
which elements are in it and how often they occur. I use variants of
this recipe. I can also collect attributes just as easily, as you can
see. If you add a stack, you can keep track of which elements occur
within other elements (for this, of course, you also have to override
the
endElement
method so you can pop the stack).
This recipe also works well as a simple example of a SAX application,
usable as the basis for any SAX application. Alternatives to SAX
include pulldom and minidom. These would be overkill for this simple job, though. For any simple processing, this is generally the case, particularly if the document you are processing is ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access