![]() | Google SRE treats reliability as engineered trade-off — SLOs define acceptable failure, error budgets cap risk-taking, and automation plus on-call culture keeps humans for judgment, not repetition. Beyer (2016) · O’Reilly Media · ISBN 9781491929124 |
|---|
Key principles
- Service level objective — Measurable reliability target — promises must be numeric, not heroic.
- Error budget policy — Allowed unreliability funds innovation — zero-failure culture kills shipping or lies about uptime.
- Toil reduction — Manual repetition is debt — automate before heroics normalize burnout.
- Monitoring and observability — Measure what users feel — dashboards serve response, not decoration.
- Blameless postmortem — Learn from incident without verdict theatre — systems fail, people operate systems.
- Release discipline — Gradual rollout and rollback — speed without reversibility is gambling.
Core science
Site Reliability Engineering documents Google-scale production practice — case studies and patterns transferable to personal operating systems when mapped carefully. On the Codex shelf pair with Observability for feedback loops, error budgets, and honest measurement.
Application
Reach for this volume when reliability, observability, or error budgets map to personal systems design — before adding tools without SLO thinking. Open Observability on the Atlas shelf for the mechanism map.
