Site Reliability Engineering (SRE)
Site reliability engineering (SRE) is a discipline, originating at Google, that applies software engineering to operations, using service level objectives, error budgets, automation and blameless reviews to keep services reliable.
Site reliability engineering (SRE) is a discipline that applies software engineering to infrastructure and operations problems in order to run reliable, scalable services. It originated at Google, where it was framed as the result of asking software engineers to design an operations function. SRE teams define reliability targets as service level objectives, automate repetitive operational work, known as toil, and share responsibility for production with development teams.
Core practices include error budgets that balance release speed against reliability, blameless post-incident reviews, capacity planning, on-call rotations with limits on operational load and progressive rollouts. Google's published SRE books made the practices widely known, and many organisations now run SRE teams or embed SRE practices in platform engineering. SRE is often described as one concrete way of implementing DevOps principles.
Adopting SRE starts with measuring what users experience, through service level indicators such as availability and latency, rather than server health alone. Reliability targets are set deliberately below 100 per cent, because perfect reliability is neither achievable nor economic. Toil reduction, observability and incident management are recurring investments, and small organisations often adopt the practices without a dedicated SRE team.
Key points
- Originated at Google and popularised through its published SRE books
- Reliability targets are expressed as service level objectives with error budgets
- Repetitive manual operational work, called toil, is systematically automated
- Blameless post-incident reviews focus on system fixes rather than individuals
Where AiVibe comes in
AiVibe Software Services delivers cloud solutions on AWS, Microsoft Azure, Google Cloud or on-premise, together with cloud security, legacy modernisation, data analytics and AI and machine learning services.
Related terms
- Service Level Objective (SLO)Cloud & AI Infrastructure
- ObservabilityCloud & AI Infrastructure
- DevOpsCloud & AI Infrastructure
- Canary ReleaseCloud & AI Infrastructure
- High Availability (HA)Cloud & AI Infrastructure
- Mean Time to Repair (MTTR)Quality, Reliability & Maintenance