Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Collector export queues: measure the failure budget before telemetry is dropped

Last updated: 1 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

An OpenTelemetry Collector exporter can queue telemetry while its destination is slow or unreachable. In-memory buffering is lost on a collector crash; a persistent queue can survive a restart but still has finite disk and retry limits. A receiver accepting a span is therefore not the same as the backend storing it. The Collector's own queue size, capacity, enqueue-failure, and send-failure metrics reveal where that guarantee ends.

Operational decision

A trace gateway receives 4,200 spans per second. Its exporter can send only 2,700 while the backend is degraded. Measure the queue in batches and the average spans per batch before translating capacity to time: a queue sized for 300,000 spans fills in about 200 seconds at a 1,500-span-per-second deficit. The expression below reports queue occupancy as a fraction for matching exporter label sets; confirm the exact metric names and labels on the installed Collector distribution. Inject a backend outage in an isolated environment, then observe queue growth, failed enqueues, retry attempts, and the last trace visible at the destination. If retention across a collector restart is required, test a persistent queue on a disk with a separate capacity alert. Restart the gateway while the queue contains known trace IDs and verify those IDs arrive once the backend returns. The acceptance record must state whether duplicates, drops, or both can occur after a timeout.

promql
otelcol_exporter_queue_size / clamp_min(otelcol_exporter_queue_capacity, 1)

Cost and verification

Memory queues compete with the collector's processors and receiver for heap. Persistent queues spend disk writes and can fail when that disk is full. Batching lowers request overhead but makes each queued item larger, so count-based queue capacity alone is not a byte budget. Scaling gateways can spread new traffic but does not recover data already dropped. Measure both queue occupancy and destination acceptance under sustained load.

Common Mistakes

  • Do not equate a successful OTLP receive with durable backend storage.
  • Do not call a disk-backed queue unlimited or assume it survives a failed disk.
  • Do not size capacity from batch count without measuring batch size and retry time.

Connected lessons

Practice and check

devops
operations
Storage details