Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Scrape staleness: separate a failed target from a missing target

Last updated: 1 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A Prometheus scrape failure and a missing target are different states. A discovered target that cannot be scraped normally reports an up sample of zero. When discovery stops reporting that target, its series can become stale and disappear from an instant query. An alert based only on up equal to zero therefore misses a target removed from discovery. The monitoring system also needs an inventory expectation or an independent user-path check; an absent series cannot prove which workload disappeared.

Operational decision

A receipt API has three intended instances. During a routing incident, one instance still appears in discovery but refuses connections; later a bad discovery label removes all instances from the scrape job. Compare target inventory with an expected deployment count before choosing the alarm. The expression below covers a failed scrape and the complete absence of recent up samples for the job, using a seven-minute window selected for this exercise. Test both faults separately. Check the scrape interval, target relabeling, and rule evaluation interval, then measure when the alert becomes active and when a synthetic receipt request fails. If only one of three targets disappears, the job-level absence branch stays false; a separate expected-instance signal or service-discovery audit is required. Do not fill missing measurements with zero in the SLO calculation without first deciding whether the service or the telemetry path failed.

promql
up{job="receipt-api"} == 0 or absent_over_time(up{job="receipt-api"}[7m])

Cost and verification

The absence window trades speed for resistance to short scrape interruptions. A very short window pages during a planned target shuffle; a long window delays detection of a monitoring blind spot. The expression scans recent up samples, while an expected-replica rule requires additional state. Check query output under both faults and verify a separate user-path signal, because green telemetry collection can coexist with a broken checkout or receipt endpoint.

Common Mistakes

  • Do not infer service health from a series that disappeared from the query.
  • Do not use a job-level absence test to claim that every individual replica is present.
  • Do not turn missing telemetry into successful requests in an SLO denominator.

Connected lessons

Practice and check

devops
operations
Storage details