Sample expensive traces without sampling away the event counts and identities needed for incident and quality analysis.
Trace sampling and coverage: know what production evidence misses
Separate counters from sampled detail
Request and failure counters should describe all traffic, while detailed traces can be sampled to control cost. A low trace rate may still help diagnose typical latency but can miss rare timeouts or feature failures. Keep a compact durable decision record where later outcome joins require one, subject to the privacy and retention policy. Restricted inference logging carries that identity; bounded metrics report the population. Do not derive an exact failure count by counting sampled traces.
Make sampling rules explicit
Head sampling decides before knowing whether a request will fail; tail sampling can retain completed error traces but requires buffering and enough time to assemble spans. Record sampling policy revision and expected probability. Ensure a trace carries model digest, policy revision and region without raw receipt contents. If errors are kept at a higher rate than successes, the sampled set is biased; comparing their raw counts does not estimate production error rate. Use complete counters for rates and sampled traces for causal investigation.
Track missing events as a first-class metric
Compare ingress counts, accepted inference counts, decision-log writes and trace samples by time bucket. The gaps should be explained: rejected requests, logging failures, sampling or delayed delivery. A telemetry outage can coincide with a model outage, so silence is not evidence of health. Outcome coverage needs its own numerator and denominator because labels can be delayed or selectively observed. A sampled trace is not a replacement for the stable decision ID in the outcome ledger.
Exercise the pipeline during failure
Drop trace export for one interval and verify serving continues while a telemetry-health alert fires. Force a model timeout and ensure a complete failure counter increments even when the trace is not sampled. Inspect tail-sampling buffer pressure and late span loss. The project compares full request counts with sampled traces and shows a missing-log incident that a trace dashboard alone would hide. State the blind spots in the on-call runbook.
Implementation
def coverage_audit(ingress, accepted, decisions, rejected):
counts = (ingress, accepted, decisions, rejected)
if any(value < 0 for value in counts):
raise ValueError("negative count")
unexplained = ingress - decisions - rejected
if accepted < decisions or unexplained < 0:
return {"state": "invalid", "gap": unexplained}
return {"state": "investigate" if unexplained else "reconciled",
"gap": unexplained}
assert coverage_audit(470, 450, 450, 20)["state"] == "reconciled"
assert coverage_audit(470, 450, 447, 20) == {"state": "investigate",
"gap": 3}
Performance and operating cost
The count comparison is O(1) time and space after aggregation. Tail sampling costs buffer memory proportional to concurrent traces and retention time; complete decision logging costs durable writes proportional to accepted requests. Counts from different windows or delayed pipelines cannot be compared directly, so align watermarks before declaring a gap.
Common Mistakes
- Using sampled trace counts as exact production failure rates.
- Keeping errors at a higher rate without reporting sampling bias.
- Treating no traces during export failure as no requests.
- Joining outcomes only to trace IDs that happen to be sampled.
Read next
- Model monitoring dimensions without metric-cardinality failure
- Project: measure receipt inference without losing the denominator
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Prediction-outcome joins: evaluate only mature, matched decisions
- Model incidents: build a release and evidence timeline before rollback
Continue the workflow: Offline edge telemetry: late events, coverage and rollback.
Continue the workflow: Runtime patch canaries: security fix without model drift.
Continue the workflow: Probe alert quality: coverage, noise and failure rehearsal.
