Executive Summary
Site Reliability Engineering (SRE) is an emerging IT service management (ITSM) framework that is essential to defending the reliability of an organization’s service by balancing two often competing demands: change/feature release velocity and site reliability. SRE methodology aligns teams on a common strategy for change management. Executives, product owners, developers, and Site Reliability Engineers agree upon a standard definition and acceptable level of reliability and decide what will happen if the organization fails to meet those standards. Organizations use these well-defined, concrete goals to set internal and external expectations for stability and to manage system changes against these specific performance metrics. The result is lower operational costs, enhanced development productivity, and increased feature release.
SRE and DevOps methodologies emphasize different metrics for measuring IT infrastructure and improving software delivery performance. The Accelerate State of DevOps Report published by DevOps Research & Assessment (DORA) identifies four key metrics for measuring software development and delivery (which some refer to as “DevOps performance”): deployment frequency, lead time for changes, time to restore service, and change failure rate. Although these metrics are important, focusing on speed and stability alone is not sufficient for organizations that deliver services and applications online.
DORA’s report also finds that elite DevOps ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access