Alertmanager groups, routes, and inhibits alerts after rules fire. Grouping reduces duplicate notifications for related alerts. Inhibition can suppress a symptom only while a designated source alert matches the same scope. A silence is a separate, time-bounded matcher used for known maintenance. Broad matchers can hide unrelated failures, and a route that sends a page to the wrong owner can make a correct detection operationally useless.
Alert routing and inhibition: suppress symptoms without silencing the cause
Operational decision
A settlement cluster loses database connectivity. Twenty instance-level failure alerts fire alongside one cluster-level database alert. Configure a reviewed policy that routes the source alert to the database owner and inhibits only instance symptoms in the same cluster while that source is active. The text fragment is a policy contract, not a complete Alertmanager configuration. In a staging alert stream, inject the source and symptoms in two clusters, then confirm the second cluster still pages. Remove the source alert and prove symptoms become notifiable again. Test grouping delay against the service's page-time objective; waiting too long for a root alert can delay a real page. For maintenance, create a narrowly scoped silence with an expiry and an owner, then verify it cannot match future clusters. Keep an independent notification path for a complete Alertmanager outage.
Settlement alert policy
Root alert: DatabaseUnavailable, owner database on-call
Symptom alert: SettlementInstanceUnavailable
Inhibit only when cluster labels are equal
Group by cluster and alert name; retain affected instance list
Maintenance silence: named owner, narrow matcher, expiry
Test: other cluster still pages; source removal restores symptom pageCost and verification
Grouping saves responder attention but may delay the first notification. Inhibition lowers noise while the source is visible; if source detection fails, symptoms must still route. A broad silence can hide a separate incident at almost no compute cost, which makes review and expiry more important than config size. Measure page latency, duplicate notifications, false suppression, and unowned routes in the staging test.
Common Mistakes
- Do not inhibit symptoms across clusters because their alert names match.
- Do not make a maintenance silence permanent or ownerless.
- Do not assume alert-rule evaluation proves the notification reached a human.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Alert design: page on impact and include a first action
- SLO burn-rate alerts: page on budget consumption, not isolated spikes
- Incident response: contain impact, then learn
- Scrape staleness: separate a failed target from a missing target
