March 2019
Beginner
490 pages
12h 40m
English
This is the code for our first spider. Save it in a file named MySpider.py under the spiders directory in your project:
from scrapy.contrib.spiders import CrawlSpider, Rulefrom scrapy.linkextractors.lxmlhtml import LxmlLinkExtractorfrom scrapy.selector import HtmlXPathSelectorfrom scrapy.item import Itemclass MySpider(CrawlSpider): name = 'example.com' allowed_domains = ['example.com'] start_urls = ['http://www.example.com'] rules = (Rule(LxmlLinkExtractor(allow=()))) def parse_item(self, response): hxs = HtmlXPathSelector(response) element = Item() return element
CrawlSpider provides a mechanism that allows you to follow the links that follow a certain pattern. Apart from the inherent attributes of the BaseSpider class, ...
Read now
Unlock full access