Two independent Prometheus servers can scrape the same targets so one server outage does not erase live visibility. Each server may send a copy of the samples to remote storage and a copy of firing alerts to Alertmanager. A remote query layer must identify replicas before deduplicating their samples; Alertmanager handles identical alert notifications through its own grouping and deduplication path. Neither behavior follows merely from deploying two replicas.
Monitoring redundancy: keep duplicate collectors without double-counting requests
Operational decision
Run two scrape servers for the receipt API on separate nodes. Give their remote samples a stable cluster identity and distinct replica identities, then configure the long-term query layer's deduplication rule explicitly. Send identical alert labels to every configured Alertmanager peer and test one scrape-server shutdown followed by one Alertmanager-peer shutdown. Query the receipt request counter before and after each fault: the service total must not double while both scrapers run, and it must not drop to zero when one stops. Compare alert receipt and notification logs with the expected page count. During an Alertmanager network partition, duplicate notifications may be preferable to missing a page; document that behavior for on-call staff. Keep both scrape targets and both alert routes under configuration review, or an apparently redundant pair may share the same failure domain.
Receipt monitoring HA contract
Scrape pair: separate nodes and independent local storage
External identity: one cluster, two distinct replica values
Remote query: tested replica deduplication rule
Alert labels: identical across scrape replicas
Alertmanager: each Prometheus targets every peer
Fault tests: one scraper down, one alert peer down, partition
Acceptance: counter neither doubles nor disappears; page still arrivesCost and verification
A second scraper roughly doubles scrape and local storage work and may duplicate remote-write traffic. Query deduplication reduces double-counting but can mask a replica with partial data if its rules are not understood. More alert peers add network and state-management work. Measure real query counts and page delivery during faults; a green HA topology diagram is not evidence.
Common Mistakes
- Do not add a replica label to alert identity if identical alerts must deduplicate.
- Do not assume remote storage removes duplicate samples without an explicit rule.
- Do not put a single load balancer in front of the only Alertmanager peer list.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Observability: join metrics, logs, and traces
- Remote-write backlog: budget the gap between local samples and long-term storage
- Alert routing and inhibition: suppress symptoms without silencing the cause
- Incident response: contain impact, then learn
