Keep operational metrics bounded while retaining detailed, restricted events for per-decision investigations.
Model monitoring dimensions without metric-cardinality failure
Choose bounded dimensions
A metric may be split by model family, approved digest, route, region and status. Each label combination creates another time series, so adding receipt ID or merchant ID can multiply storage until monitoring fails during the incident it should reveal. Use controlled vocabularies and cap version retention. Keep request identifiers in restricted logs or traces instead. Decision logs support exact joins; alerts need only dimensions that lead to a different owner or action.
Match metric type to question
Count requests, failures and fallback as counters; observe latency distributions with buckets suited to the service deadline; export active in-flight work as a gauge. Track numerator and denominator separately so a rate cannot hide changing traffic volume. A deployment might leave an old digest label active for a short rollback window, but it should not create a fresh metric name for every run ID. Tail latency requires a distribution, not an average of worker times.
Budget the series before adding labels
Estimate the product of label values for each metric family across replicas and retained versions. A few bounded dimensions can still create thousands of series when combined. Put a ceiling in review and compare planned series growth with monitoring capacity. If a question needs arbitrary merchant-level cuts, compute it from access-controlled event data rather than live time-series labels. A single “other” bucket can contain unexpected values while preserving counts; investigate the bucket rather than silently dropping it.
Test observability under change
Roll out a new digest and confirm old series age out on schedule. Force a fallback and verify both total request and fallback counters move. Scrape during a burst and check that the telemetry pipeline does not add material serving latency. The project injects a receipt-ID label proposal and rejects it before production, then tests a bounded alternative that still identifies model-specific failure. Monitoring cost belongs in capacity planning, not just the dashboard backlog.
Implementation
def estimate_series(label_values, replicas):
if replicas < 0 or any(count < 0 for count in label_values.values()):
raise ValueError("negative series dimension")
total = replicas
for count in label_values.values():
total *= count
return total
bounded = {"model": 3, "region": 2, "route": 3, "status": 4}
assert estimate_series(bounded, 5) == 360
assert estimate_series({**bounded, "receipt_id": 47000}, 5) == 16920000
Performance and operating cost
The estimate is O(d) time and O(1) extra space for d dimensions. Real series count depends on which combinations actually occur, but the product is a useful upper bound. Each extra series consumes memory, storage, network and query work. The 47,000 receipt IDs are fictional; the calculation shows why an unbounded identifier is unsuitable as a metric label.
Common Mistakes
- Putting decision IDs or customer IDs in metric labels.
- Exporting only a rate without the numerator and denominator.
- Averaging worker latency while queue time is omitted.
- Keeping all historical digest label sets active forever.
Read next
- Trace sampling and coverage: know what production evidence misses
- Project: measure receipt inference without losing the denominator
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Model alerts: page on customer symptoms with a named owner
- Inference latency budgets: measure queue, feature and model time
