Chapter 17. Avoiding Scraping Traps
Few things are more frustrating than scraping a site, viewing the output, and not seeing the data that’s so clearly visible in your browser. Or submitting a form that should be perfectly fine but gets denied by the web server. Or getting your IP address blocked by a site for unknown reasons.
These are some of the most difficult bugs to solve, not only because they can be so unexpected (a script that works just fine on one site might not work at all on another, seemingly identical, site), but because they purposefully don’t have any telltale error messages or stack traces to use. You’ve been identified as a bot, rejected, and you don’t know why.
In this book, I’ve written about a lot of ways to do tricky things on websites, including submitting forms, extracting and cleaning difficult data, and executing JavaScript. This chapter is a bit of a catchall in that the techniques stem from a wide variety of subjects. However, they all have something in common: they are meant to overcome an obstacle put in place for the sole purpose of preventing automated scraping of a site.
Regardless of how immediately useful this information is to you at the moment, I highly recommend you at least skim this chapter. You never know when it might help you solve a difficult bug or prevent a problem altogether.
A Note on Ethics
In the first few chapters of this book, I discussed the legal gray area that web scraping inhabits, as well as some of the ethical and legal ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access