An OpenTelemetry Collector exporter can queue telemetry while its destination is slow or unreachable. In-memory buffering is lost on a collector crash; a persistent queue can survive a restart but still has finite disk and retry limits. A receiver accepting a span is therefore not the same as the backend storing it. The Collector's own queue size, capacity, enqueue-failure, and send-failure metrics reveal where that guarantee ends.
Collector export queues: measure the failure budget before telemetry is dropped
Operational decision
A trace gateway receives 4,200 spans per second. Its exporter can send only 2,700 while the backend is degraded. Measure the queue in batches and the average spans per batch before translating capacity to time: a queue sized for 300,000 spans fills in about 200 seconds at a 1,500-span-per-second deficit. The expression below reports queue occupancy as a fraction for matching exporter label sets; confirm the exact metric names and labels on the installed Collector distribution. Inject a backend outage in an isolated environment, then observe queue growth, failed enqueues, retry attempts, and the last trace visible at the destination. If retention across a collector restart is required, test a persistent queue on a disk with a separate capacity alert. Restart the gateway while the queue contains known trace IDs and verify those IDs arrive once the backend returns. The acceptance record must state whether duplicates, drops, or both can occur after a timeout.
otelcol_exporter_queue_size / clamp_min(otelcol_exporter_queue_capacity, 1)Cost and verification
Memory queues compete with the collector's processors and receiver for heap. Persistent queues spend disk writes and can fail when that disk is full. Batching lowers request overhead but makes each queued item larger, so count-based queue capacity alone is not a byte budget. Scaling gateways can spread new traffic but does not recover data already dropped. Measure both queue occupancy and destination acceptance under sustained load.
Common Mistakes
- Do not equate a successful OTLP receive with durable backend storage.
- Do not call a disk-backed queue unlimited or assume it survives a failed disk.
- Do not size capacity from batch count without measuring batch size and retry time.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Distributed traces: preserve context without leaking data
- Trace sampling: retain useful failures without flooding storage
- Observability: join metrics, logs, and traces
- Remote-write backlog: budget the gap between local samples and long-term storage
