Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Tail sampling at scale: keep every span of a trace with one decision maker

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Tail sampling waits for enough of a trace to arrive before deciding whether to keep it. That allows a policy to retain error or slow traces, but the collector must hold spans and decision state during the wait. A round-robin balancer can split one trace across collectors; each replica then sees an incomplete history and may reach a different decision. A trace-aware routing tier keeps spans with the same trace ID together, while late spans and replica changes still need an explicit policy.

Operational decision

A payment service emits traces across an API, a queue consumer, and a ledger writer. Place a routing tier before two sampling gateways and send all spans for a given trace ID to the same gateway. Generate a slow trace, an error trace, and a normal trace; verify that each retained trace contains the expected cross-service spans. Now restart one gateway during the decision wait and record which traces become partial or are lost. Keep the decision window long enough for typical queue delay, but do not assume it can cover an unbounded async workflow. Measure memory consumed by traces awaiting a decision, the rate of spans arriving after a decision, and routing imbalance. Use span links for separate asynchronous work rather than forcing unrelated operations under one trace ID. If the routing topology changes, retest decisions before using sampled traces as incident evidence.

Output
Tail-sampling acceptance record
Routing key: trace ID, stable across gateway tier
Policy: keep error and slow traces, sample normal traces
Decision wait: compared with measured async span delay
Restart drill: partial-trace rate and lost-decision count
Late span rule: recorded and tested
Acceptance: expected spans present for retained test traces

Cost and verification

A longer decision wait raises memory use and delays trace availability. More sampling gateways add routing and state-transfer risk; they do not automatically improve a trace-aware decision. Keeping every slow trace can produce a surge when the whole service slows, so capacity needs a degraded-mode policy. Compare full-trace completeness and collector memory under both ordinary and incident load.

Common Mistakes

  • Do not round-robin individual spans into independent tail samplers.
  • Do not assume a retained root span proves the whole async operation was captured.
  • Do not choose a decision wait without measuring late-span behavior.

Connected lessons

Practice and check

devops
operations
Storage details