Problem details
We often forget that websites are not just used by humans. A significant percentage of web traffic comes from other programs such as crawlers, bots, or scrapers. Sometimes, you will need to write such programs yourself to extract information from another website.
Generally, pages designed for human consumption are cumbersome for mechanical extraction. HTML pages have information surrounded by markup, requiring extensive cleanup. Sometimes, information will be scattered, needing extensive data collation and transformation.
A machine interface would be ideal in such situations. You cannot only reduce the hassle of extracting information, but also enable the creation of mashups. The longevity of an application will be greatly ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access