Skip to Content
Site Reliability Engineering
book

Site Reliability Engineering

by Niall Richard Murphy, Betsy Beyer, Chris Jones, Jennifer Petoff
April 2016
Intermediate to advanced
552 pages
15h 44m
English
O'Reilly Media, Inc.
Audiobook available
Content preview from Site Reliability Engineering

Appendix D. Example Postmortem

Shakespeare Sonnet++ Postmortem (incident #465)

Date: 2015-10-21

Authors: jennifer, martym, agoogler

Status: Complete, action items in progress

Summary: Shakespeare Search down for 66 minutes during period of very high interest in Shakespeare due to discovery of a new sonnet.

Impact:1 Estimated 1.21B queries lost, no revenue impact.

Root Causes:2 Cascading failure due to combination of exceptionally high load and a resource leak when searches failed due to terms not being in the Shakespeare corpus. The newly discovered sonnet used a word that had never before appeared in one of Shakespeare’s works, which happened to be the term users searched for. Under normal circumstances, the rate of task failures due to resource leaks is low enough to be unnoticed.

Trigger: Latent bug triggered by sudden increase in traffic.

Resolution: Directed traffic to sacrificial cluster and added 10x capacity to mitigate cascading failure. Updated index deployed, resolving interaction with latent bug. Maintaining extra capacity until surge in public interest in new sonnet passes. Resource leak identified and fix deployed.

Detection: Borgmon detected high level of HTTP 500s and paged on-call.

Action Items:3

Action Item Type Owner Bug

Update playbook with instructions for responding to cascading failure

mitigate

jennifer

n/a DONE

Use flux capacitor to balance load between clusters

prevent

martym

Bug 5554823 TODO

Schedule cascading failure test during next DiRT ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Start your free trial

You might also like

Site Reliability Engineering Fundamentals

Site Reliability Engineering Fundamentals

Emil Stolarsky, Jaime Woo
Observability Engineering

Observability Engineering

Charity Majors, Liz Fong-Jones, George Miranda
The Site Reliability Workbook

The Site Reliability Workbook

Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara, Stephen Thorne
AI Engineering

AI Engineering

Chip Huyen

Publisher Resources

ISBN: 9781491929117Errata Page