What is considered the best practice when working with alerting notifications?
The Prometheus alerting philosophy emphasizes signal over noise --- meaning alerts should focus only on actionable and user-impacting issues. The best practice is to alert on symptoms that indicate potential or actual user-visible problems, not on every internal metric anomaly.
This approach reduces alert fatigue, avoids desensitizing operators, and ensures high-priority alerts get the attention they deserve. For example, alerting on ''service unavailable'' or ''latency exceeding SLO'' is more effective than alerting on ''CPU above 80%'' or ''disk usage increasing,'' which may not directly affect users.
Option B correctly reflects this principle: keep alerts meaningful, few, and symptom-based. The other options contradict core best practices by promoting excessive or equal-weight alerting, which can overwhelm operations teams.
Verified from Prometheus documentation -- Alerting Best Practices, Alertmanager Design Philosophy, and Prometheus Monitoring and Reliability Engineering Principles.
Amber
5 months agoMurray
5 months agoLauran
6 months agoBok
6 months agoMyra
6 months agoNathalie
6 months agoOctavio
6 months agoCristina
7 months agoSabra
7 months agoMarion
7 months agoJulieta
7 months agoCecilia
8 months agoBlondell
8 months agoMichal
8 months agoCarlton
8 months agoJess
8 months agoKayleigh
8 months agoYasuko
9 months agoAnnamae
9 months agoJose
9 months agoJesus
9 months agoMelina
9 months agoJin
10 months agoAlba
10 months agoXuan
10 months agoEdmond
10 months agoYvette
4 months agoHershel
4 months agoMan
5 months agoJulian
5 months agoGail
5 months ago