July 2017
Beginner to intermediate
312 pages
7h 27m
English
A scraping framework has two main independent components. The spider crawls the site and the pipeline processes the scraped data, as shown in this diagram:

These two components are independent because the spider is dependent on the format and structure of the website whereas the pipeline is dependent on the structure of the persisted data.
The spider works as follows: it takes a URL as the entry point (for example, the landing page), extracts all the links present on the page, and crawls to the next page. On the next page, the spider will repeat the crawling process until it reaches a predefined level of depth. On a forum, usually ...
Read now
Unlock full access