Trace sampling decides which distributed requests are retained for analysis. Head sampling chooses near the start of a trace, which is cheap but cannot know whether a later span fails. Tail sampling waits to inspect much of the trace, enabling error or latency criteria at the cost of buffering, state, and a later decision. Neither method repairs missing propagation between services.
Trace sampling: retain useful failures without flooding storage
Operational decision
A parcel checkout has high traffic but rare payment timeouts. Keep a low, consistent baseline of successful traces and a higher share of error and slow traces. The text policy is a design record; implement it in the chosen collector and SDK with supported settings. Send all spans of a trace to the same decision point, otherwise a tail sampler may see only a fragment. Size the wait period to cover the expected trace duration and set a memory limit for bursts. During a controlled timeout, verify the collector retains the trace containing the downstream error and that trace IDs connect to logs. Measure refused or dropped spans at both SDK and collector. If a head sampler already discarded the trace, a later tail rule cannot recover it. Keep secrets and personal data out of attributes regardless of sampling rate.
Parcel trace policy
Baseline: consistent low-rate successful traces
Priority: errors and slow checkout paths
Tail wait: covers measured request duration
Routing: all spans of one trace reach one decision point
Guard: bounded collector memory and drop alert
Proof: injected payment timeout appears with correlated logsCost and verification
Head sampling uses less collector memory and exports less data early; tail sampling provides richer selection but must buffer spans while a decision is pending. More retained traces increase storage and transfer cost. A policy optimized only for errors can miss slow successful requests or a broken sampler. Track accepted, sampled, and dropped spans, and compare with independently measured request failures. Treat sampling as a data-volume control, not a substitute for error-rate metrics.
Common Mistakes
- Do not expect tail sampling to recover traces dropped at the head.
- Do not split one trace across independent tail samplers without consistent routing.
- Do not put private payloads into spans because most traces are sampled out.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Distributed traces: preserve context without leaking data
- Metric cardinality: keep observability usable during a surge
- Log pipelines: preserve incident evidence without ingesting secrets
- Incident response: contain impact, then learn
