Preface
Two previous O’Reilly books from Google—Site Reliability Engineering by Betsy Beyer, Chris Jones, Niall Richard Murphy, and Jennifer Petoff (Eds.) and The Site Reliability Workbook by Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara, and Stephen Thorne (Eds.)—demonstrated how and why a commitment to the entire service lifecycle enables organizations to successfully build, deploy, monitor, and maintain software systems.
This report is designed to build on the foundation of those books and delve a little deeper into the challenges of adopting site reliability engineering (SRE) in large and complex organizations (which we refer to as enterprises). Despite the popularity of SRE over the past few years, we have feedback from numerous enterprises that there is a gap between the enthusiasm for SRE and the level of adoption.
We think this is an important gap to close because reliability is increasingly a major differentiator for enterprises. The pace and scale of technology change triggered by both cloud adoption and the COVID-19 pandemic often requires different techniques to handle this increased complexity.
These topics will be of more interest to you if you are involved in (or depend on) reliability for production systems and need to know more about SRE adoption. This includes executive and leadership roles but also individual contributors (cloud architects, site reliability engineers [SREs], platform developers, etc.) Regardless of role, if you design, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access