Slos
4 pages
-
Availability
concept
Uptime/downtime definition; the nines table (90%–99.999%); techniques for high availability; dependency chaining effects; design-for-production philosophy; ROI of availability investment; MTBF/MTTR/RPO/RTO measurement framework; tyranny of the nines antipattern; Allspaw: MTTR > MTBF; SRE: 100% is always the wrong target
-
Error Budgets
concept
error budget = 1 − SLO target; resolves dev/ops conflict by aligning incentives; budget exhaustion triggers release freeze; burn rate alerting; 100% is wrong target argument
-
Monitoring
concept
Black-box vs white-box monitoring; metrics and pre-aggregation; SLIs/SLOs; burn rate alerting; three-layer ML monitoring taxonomy (golden signals/generic ML signals/domain-specific quality); four actuals cases; drift detection (PSI, KL divergence, Wasserstein); ML SLOs and privacy in monitoring
-
Site Reliability Engineering
source
*Site Reliability Engineering* — Beyer, Jones, Petoff, Murphy