Site Reliability Engineering Essential Training
4h 10mIntermediate2025-07-01
Authors

Pearson

Karun Subramanian
Course details
Unlock the power of Site Reliability Engineering (SRE) with this comprehensive video course. SRE is a critical discipline that combines software engineering with IT operations to ensure high system reliability, scalability, and performance. This course provides a deep dive into the core principles and practices of SRE, equipping you with the tools to build reliable systems and improve operational efficiency.
Learn key SRE concepts, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, with practical examples that help you apply these principles to your own organization. The course also addresses crucial aspects of incident management, such as managing on-call duties, running war rooms for critical incidents, and conducting blameless postmortems to learn from failures. Additionally, discover release management strategies that minimize user impact during deployments, monitor your CI/CD pipeline, and ensure progressive rollouts.
Learning objectives
Set a strong foundation by implementing core Site Reliability Engineering (SRE) principles to ensure system reliability and performance.
Build and optimize a robust monitoring and observability system using essential telemetry data such as logs, metrics, and traces.
Monitor system health effectively through observability platforms to maintain optimal system performance.
Apply Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to improve system reliability and performance.
Manage incidents effectively, run war rooms for critical situations, and conduct blameless postmortems to learn from failures.
Design reliable system architectures, including load balancing, auto-scaling, and implementing the CAP theorem for system resilience.
Learn key SRE concepts, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, with practical examples that help you apply these principles to your own organization. The course also addresses crucial aspects of incident management, such as managing on-call duties, running war rooms for critical incidents, and conducting blameless postmortems to learn from failures. Additionally, discover release management strategies that minimize user impact during deployments, monitor your CI/CD pipeline, and ensure progressive rollouts.
Learning objectives
Set a strong foundation by implementing core Site Reliability Engineering (SRE) principles to ensure system reliability and performance.
Build and optimize a robust monitoring and observability system using essential telemetry data such as logs, metrics, and traces.
Monitor system health effectively through observability platforms to maintain optimal system performance.
Apply Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to improve system reliability and performance.
Manage incidents effectively, run war rooms for critical situations, and conduct blameless postmortems to learn from failures.
Design reliable system architectures, including load balancing, auto-scaling, and implementing the CAP theorem for system resilience.
Skills covered
DevOps FoundationsServer AdministrationDevOpsNetwork and System AdministrationOne-Off
Concepts
0. Introduction
- 01 - Introduction
1. Introduction to Site Reliability Engineering
- 02 - Learning objectives
- 03 - What is site reliability engineering
- 04 - Core tenets of SRE
- 05 - Benefits of SRE
- 06 - DevOps vs. SRE vs. platform engineering
- 07 - A typical day of an SRE
2. Observability
- 08 - Learning objectives
- 09 - What to monitor
- 10 - Logs, metrics, and traces
- 11 - The four golden signals
- 12 - Observability platforms
- 13 - Demo - Monitoring using Splunk
3. SLO, SLI, and SLA
- 14 - Learning objectives
- 15 - Service-level objectives (SLO)
- 16 - Service-level indicators (SLI) and service-level agreements (SLA)
- 17 - Implementing SLOs - Real-world examples
- 18 - Using error budgets
- 19 - Demo - SLO SLI
4. Incident Management, SRE Style
- 20 - Learning objectives
- 21 - Managed vs. unmanaged incidents
- 22 - Running war rooms
- 23 - Conducting blameless postmortems
- 24 - Using postmortem templates
- 25 - Being on call
5. Reliable System Architectures
- 26 - Learning objectives
- 27 - Load balancing
- 28 - Handling failures
- 29 - CAP theorem and its implementation
- 30 - Auto scaling
6. Release Management
- 31 - Learning objectives
- 32 - Progressive rollout
- 33 - Minimizing user impact during releases
- 34 - Monitoring the CI CD pipeline
- 35 - Rolling back changes
7. Implementing SRE
- 36 - Learning objectives
- 37 - Four ways of implementing in your organization
- 38 - Benefits of a central SRE team
- 39 - Production readiness review
8. Course Conclusion and Next Steps
- 40 - Learning objectives
- 41 - Course summary
- 42 - Next steps
Conclusion
- 43 - Wrapping up