Module: XML Lexing (Shallow Parsing)
Credit: Paul Prescod
It’s not uncommon to want to work with the form of an XML document rather than with the structural information it contains (e.g., to change a bunch of entity references or element names). The XML may be slightly incorrect, enough to choke a traditional parser. In such cases, you need an XML lexer, also known as a shallow parser.
You might be tempted to hack together a regular expression or two to do some simple parsing of XML (or other structured text format), rather than using the appropriate library module. Don’t—it’s not a trivial task to get the regular expressions right! However, the hard work has already been done for you in Example 12-1, which contains already-debugged regular expressions and supporting functions that you can use for shallow-parsing tasks on XML data (or, more importantly, on data that is almost, but not quite, correct XML, so that a real XML parser seizes up with error diagnostics when you try to parse your data with it).
A traditional XML parser does a few tasks:
It breaks up the stream of text into logical components (tags, text, processing instructions, etc.).
It ensures that these components comply with the XML specification.
It throws away extra characters and reports the significant data. For instance, it would report tag names but not the less-than and greater-than signs around them.
The shallow parser in Example 12-1 performs only the first task. It breaks up the document and presumes that you ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access