A Prometheus scrape failure and a missing target are different states. A discovered target that cannot be scraped normally reports an up sample of zero. When discovery stops reporting that target, its series can become stale and disappear from an instant query. An alert based only on up equal to zero therefore misses a target removed from discovery. The monitoring system also needs an inventory expectation or an independent user-path check; an absent series cannot prove which workload disappeared.
Scrape staleness: separate a failed target from a missing target
Operational decision
A receipt API has three intended instances. During a routing incident, one instance still appears in discovery but refuses connections; later a bad discovery label removes all instances from the scrape job. Compare target inventory with an expected deployment count before choosing the alarm. The expression below covers a failed scrape and the complete absence of recent up samples for the job, using a seven-minute window selected for this exercise. Test both faults separately. Check the scrape interval, target relabeling, and rule evaluation interval, then measure when the alert becomes active and when a synthetic receipt request fails. If only one of three targets disappears, the job-level absence branch stays false; a separate expected-instance signal or service-discovery audit is required. Do not fill missing measurements with zero in the SLO calculation without first deciding whether the service or the telemetry path failed.
up{job="receipt-api"} == 0 or absent_over_time(up{job="receipt-api"}[7m])Cost and verification
The absence window trades speed for resistance to short scrape interruptions. A very short window pages during a planned target shuffle; a long window delays detection of a monitoring blind spot. The expression scans recent up samples, while an expected-replica rule requires additional state. Check query output under both faults and verify a separate user-path signal, because green telemetry collection can coexist with a broken checkout or receipt endpoint.
Common Mistakes
- Do not infer service health from a series that disappeared from the query.
- Do not use a job-level absence test to claim that every individual replica is present.
- Do not turn missing telemetry into successful requests in an SLO denominator.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Observability: join metrics, logs, and traces
- Synthetic transactions: measure the route a user actually takes
- Alert design: page on impact and include a first action
- SLO burn-rate alerts: page on budget consumption, not isolated spikes
