Normalizing an XML Document
Credit: David Ascher, Paul Prescod
Problem
You want to compare two
different XML documents using standard tools such as
diff.
Solution
Normalize each XML document using the following recipe, then use a whitespace-insensitive diff tool:
from xml.dom import minidom dom = minidom.parse(input) dom.writexml(open(outputfname, "w"))
Discussion
Different editing tools munge XML differently. Some, like text editors, make no modification that is not explicitly done by the user. Others, such as XML-specific editors, sometimes change the order of attributes or automatically indent elements to facilitate the reading of raw XML. There are reasons for each approach, but unfortunately, the two approaches can lead to confusing differences—for example, if one author uses a plain editor while another uses a fancy XML editor, and a third person is in charge of merging the two sets of changes. In such cases, one should use an XML-difference engine. Typically, however, such tools are not easy to come by. Most are written in Java and don’t deal well with large XML documents (performing tree-diffs efficiently is a hard problem!).
Luckily, combinations of small steps can solve the problem nicely. First, normalize each XML document, then use a standard line-oriented diff tool to compare the normalized outputs. This recipe is a simple XML normalizer. All it does is parse the XML into a Document Object Model (DOM) and write it out. In the process, elements with no children are written ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access