On Call
4 pages
-
Incident Management
concept
Hypothetico-deductive troubleshooting; ICS roles; "only Ops modifies"; blameless postmortems; outage tracking; ML incident response (harder detection, broader scope, fuzzy timeline, RPO/RTO for ML)
-
Monitoring
concept
Black-box vs white-box monitoring; metrics and pre-aggregation; SLIs/SLOs; burn rate alerting; three-layer ML monitoring taxonomy (golden signals/generic ML signals/domain-specific quality); four actuals cases; drift detection (PSI, KL divergence, Wasserstein); ML SLOs and privacy in monitoring
-
Site Reliability Engineering
concept
SRE as discipline: dev/ops conflict; error budgets; toil cap (50%); SLO-driven alerting; blameless postmortems; SRE vs DevOps distinction; applicability outside Google
-
Site Reliability Engineering
source
*Site Reliability Engineering* — Beyer, Jones, Petoff, Murphy