Turning SAX Events into Data Structures
As described earlier, one of the great strengths of SAX is that it lets applications use appropriate data structures, instead of forcing the use of generic data structures. In Section 3.5.2 in Chapter 3, we looked at the problem of producing SAX events from data structures. Here we look at the reverse process: producing data structures from SAX events. This is a process that most SAX applications handle to one degree or another. One of the most traditional names for this process is unmarshaling; it’s also sometimes called deserializing. (I tend to avoid using the latter term with Java except when talking about RMI.)
We’ll first look at how to turn SAX into generic DOM (and DOM-like) data structures. If you’re working with such data structures, you may find it’s advantageous to build them using SAX. With SAX, you can easily discard data you don’t need, filtering it out so you don’t need to pay its costs. Afterward we’ll look briefly at some of the concerns associated with working with data structures that are more specialized to your application.
SAX-to-DOM Consumers
It’s easy to turn a SAX event stream into a complete DOM document tree, or into a DOM-like data structure such as DOM4J or JDOM. Most open source DOM parsers build those data structures directly from SAX event streams. (Xerces has the only such DOM I know that doesn’t work that way.) Building a DOM document from a SAX2 event stream requires implementing all four event consumer ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access