Getting a Good Night's Sleep
BCP is not about planning for unforeseen disasters; it's about preparing for the day-to-day, week-to-week, or even year-to-year random blowups that make the lives of web ops engineers hectic. If you can plan ahead, solve for the big problems, and exercise your failovers regularly as part of your day job, any failure in any part of your platform will become an easily handled event, rather than a crisis.
Remember the story I told at the beginning of the chapter, about volumes on live mail farms being deleted, because some yahoo was practicing
vol destroy? Well, just a few months back, I received the following page (paraphrased): "Sev 1 outage: 5 Flickr Photo volumes accidentally destroyed." Hmm, "accidentally destroyed"; that sounds familiar, I thought. This time it was our top storage expert, trying to repurpose an old retired filer, instead of a newbie practicing on live farms. The cause was the same, however: issuing the right command on the wrong machine. By this point, we had a solid BCP plan, and the data was already mirrored across the country. We just flipped a switch in DNS, millions of users' photos magically reappeared, and we went back to work.
Mistakes happen, both to new employees and to the most seasoned. Expect these mistakes, and plan for them; we are all human, after all. By planning for human error, machine error, and infrastructure error, you can recover from pretty much anything.
If you are going to treat BCP as a once-in-a-blue-moon ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access