January 2019
Beginner
132 pages
3h 14m
English
Contrary to whitelists, blacklists define websites where your scraper should definitely not venture. Sites that you will want to include here may be places that you know do not contain any relevant information, or you are just not interested in their content. You might also temporarily blacklist sites that are experiencing performance issues, such as a high number of 5XX errors, as discussed in Chapter 2, The Request/Response Cycle. You can match your link URLs to their hostname in the same way as in the preceding example.
The only change that is required is to modify the last if block, shown as follows, so it runs only if doesMatch is false:
if !doesMatch {// Continue scraping …}
Read now
Unlock full access