Prometheus remote write reads local write-ahead-log data and sends batches through per-destination queues. When a remote endpoint slows, pending samples and retry work grow. Queue size and shard count consume memory; more shards can also overwhelm the receiver. Local scraping can remain healthy while the long-term store falls behind, so a green target dashboard does not prove that remote queries contain the same recent history.
Remote-write backlog: budget the gap between local samples and long-term storage
Operational decision
A metrics archive for settlement services stops accepting writes for a maintenance hour. Record the incoming sample rate, pending-sample trend, remote send rate, WAL disk headroom, and time of the oldest missing sample. The fragment names the pending-sample signal; filter it by destination labels when a deployment has several remotes. During the outage, keep local Prometheus available for immediate triage, and avoid increasing queue capacity or shard count until memory and receiver capacity are measured. Restore the destination, then verify pending samples drain and compare a known counter increase in both local and remote queries for the same bounded interval. A successful HTTP response from the remote endpoint is insufficient if older samples were already lost to retention or rejected as out of order. Document the longest tolerable endpoint outage for the installed Prometheus version and configured WAL retention; do not copy a generic duration into an operational promise.
prometheus_remote_storage_samples_pendingCost and verification
A larger queue can absorb bursts but raises memory use; more parallel requests raise CPU and network cost on both sides. Recovery time depends on the difference between remote drain throughput and the new incoming sample rate, not only the backlog size. If the remote can send 18,000 samples per second while 15,000 arrive, a 5.4-million-sample backlog needs about thirty minutes to drain, ignoring retries and batching overhead. Prove the missing interval through a source-versus-remote query before closing the incident.
Common Mistakes
- Do not treat scrape success as proof that long-term storage received the samples.
- Do not raise shard count without checking receiver saturation and Prometheus memory.
- Do not call the gap recovered merely because the pending gauge reached zero.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Observability: join metrics, logs, and traces
- Metric cardinality: keep observability usable during a surge
- Alert design: page on impact and include a first action
- Capacity and load tests: identify the next bottleneck
