Build bounded metrics, sampled traces and a coverage ledger that exposes missing decision records during an incident.
Project: measure receipt inference without losing the denominator
Design the telemetry contract
For the receipt endpoint, keep full counters for ingress, accepted requests, rejected requests, decisions, fallback and errors. Observe end-to-end latency by bounded route and digest. Permit model, region, route and status labels; reject receipt ID, merchant key and free-form error text as metric labels. Put exact decision identity in restricted durable logs. The series budget calculates the largest expected metric family before any instrument is deployed.
Set trace policy and expected coverage
Sample a fixed share of ordinary successful requests and retain error traces when practical. Annotate each trace with the sampling policy revision and approved model identity. Keep full failure counters regardless of trace sampling. Write a compact decision event for accepted requests to support later outcome joins. Sampling rules must explain why 47 traces from 470 requests are not 47 total decisions or an unbiased quality sample.
Inject a telemetry failure
Send a deterministic load with valid, invalid and fallback receipts. Then drop three decision-log writes while allowing inference to continue. The ingress, rejection and decision counts should show a gap of three after their windows align. Raise a telemetry-health incident and retain the affected decision IDs from the gateway when available. Do not fill the gap by fabricating model outputs. The incident timeline should separate serving health from evidence health.
Deliver an operating record
Report configured labels, upper-bound series count, full request counts, trace sampling rate, decision-log coverage, delayed delivery window and alert owner. Run one burst with the candidate telemetry enabled and compare serving tail latency. A monitoring change that harms inference latency is not harmless. Link future quality analysis to mature outcome joins and flag the missing records so no analyst silently treats them as negative outcomes.
Implementation
def evidence_health(ingress, rejected, decisions, maximum_gap=0):
if min(ingress, rejected, decisions) < 0 or rejected + decisions > ingress:
raise ValueError("invalid evidence counts")
gap = ingress - rejected - decisions
return {"gap": gap, "state": "alert" if gap > maximum_gap else "healthy"}
assert evidence_health(470, 20, 450) == {"gap": 0, "state": "healthy"}
assert evidence_health(470, 20, 447) == {"gap": 3, "state": "alert"}
Performance and operating cost
The health calculation is O(1) time and space. Full counters are cheap, but durable decision events and trace buffers scale with traffic. A zero-gap rule may alert during normal delivery delay; compare aligned event-time windows and set a measured grace interval before paging. The deterministic counts demonstrate the contract, not a recommended production threshold.
Common Mistakes
- Adding receipt ID to a metric because it helps one dashboard query.
- Sampling decision logs required for outcome joins.
- Declaring a gap before delayed log delivery reaches its watermark.
- Assuming serving is healthy because telemetry is silent.
Read next
- Model monitoring dimensions without metric-cardinality failure
- Trace sampling and coverage: know what production evidence misses
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Model incidents: build a release and evidence timeline before rollback
- Project: operate receipt scoring with a deadline and overload path
Continue the workflow: Async result ledgers: reconcile output, failure and notification.
