A metric contract fixes the numerator, denominator, grain, time basis, filters and source version so different reports compute the same number.
Semantic metric contracts and reconciliation
Define the measure as an operation
A payment success rate could divide successful capture attempts by eligible attempts, distinct orders, or unique customers. Those are different metrics. State the event unit, which statuses count, whether retries are separate attempts, how reversals appear and which event-time window applies. Denominator design should be versioned with the same care as a source schema.
Preserve components
Store or expose the additive numerator and denominator; compute the ratio after aggregation. If one region has 3 successes from 4 attempts and another has 27 from 36, the combined rate is 30 from 40, not the unweighted mean of 75% and 75% by coincidence. Unequal rates reveal the error immediately. Do not add or average published percentages without their weights.
Choose one clock
An event-time daily rate may change when late payment responses arrive. Define a cutoff and correction window. A processing-time report answers a different question. Attach the source snapshot and metric definition version to exported results so an analyst can reproduce the observed value. Late-arrival policy controls restatement.
Reconcile each boundary
Count source attempts, excluded tests, eligible attempts, successful attempts and published metric rows. Require the categories to balance at their stated grain. Compare warehouse components with the source ledger for the same interval and snapshot. If a dimension repair reassigns a region, total attempts should stay fixed while regional buckets change.
Release a contract test
Use a fixture containing an approved capture, a declined capture, a retry, a reversed capture and an ineligible test event. Assert both component counts and final ratio. Fail publication when eligibility or join cardinality changes unexpectedly. The pipeline SLO should require a usable metric, not merely a completed query.
Implementation
payment_attempts = [
{"id": "pay-47", "eligible": True, "success": True},
{"id": "pay-48", "eligible": True, "success": False},
{"id": "pay-49", "eligible": True, "success": True},
{"id": "test-50", "eligible": False, "success": True},
]
def success_components(attempts):
if len({row["id"] for row in attempts}) != len(attempts):
raise ValueError("duplicate attempt identity")
eligible = [row for row in attempts if row["eligible"]]
return sum(row["success"] for row in eligible), len(eligible)
successful, eligible = success_components(payment_attempts)
assert (successful, eligible) == (2, 3)
def success_rate(attempts):
numerator, denominator = success_components(attempts)
return None if denominator == 0 else numerator / denominator
assert success_rate(payment_attempts) == 2 / 3
assert success_rate([]) is NonePerformance and operating cost
Scanning N attempts takes O(N) expected time and O(N) memory for identity validation. The ratio calculation is O(1), but computing it from an incorrectly joined or stale cohort is still wrong. Retaining component counts and versions adds modest storage and major audit value.
Common Mistakes
- Do not average regional rates without eligible-attempt weights.
- Do not substitute distinct orders for capture attempts without a definition change.
- Do not publish a metric without its cutoff and source version.
Read next
- Fact grain and measure additivity
- Conformed dimensions and surrogate keys
- Late-arriving dimensions and fact repair
- Project: release a reconciled revenue mart
- Metric denominators and cohorts: make a rate reproducible
Continue the workflow: Cross-engine SQL semantic differences.
Continue the workflow: Project: migrate a dashboard with shadow reads.
