Choose the decision that preserves a trustworthy user-impact signal when the telemetry system itself is degraded. A green dashboard is not evidence if collection or export has a gap.
Review these lessons
- Scrape staleness: separate a failed target from a missing target
- Latency histograms: put a bucket at the actual SLO threshold
- Remote-write backlog: budget the gap between local samples and long-term storage
- Collector export queues: measure the failure budget before telemetry is dropped
- Tail sampling at scale: keep every span of a trace with one decision maker
- Alert routing and inhibition: suppress symptoms without silencing the cause
- Monitoring redundancy: keep duplicate collectors without double-counting requests
- Telemetry redaction: remove sensitive fields before an exporter or sampler sees them
Other checks
Common Mistakes
- Do not equate local collection with remote acceptance.
- Do not silence every symptom across unrelated clusters.
