A shared feature name does not guarantee that training and serving compute the same value at the same decision time.
Feature parity: compare historical replay with online observations
Sample real decisions
Log entity ID, decision timestamp, online feature value, definition version, snapshot ID and missing state for a privacy-approved sample. Reconstruct the same decision from the historical store using its availability cutoff. Compare both value and null status. The historical join must use the actual decision time, not the day the audit runs.
Classify mismatches
Differences can come from late input, transform versions, entity-key normalization, rounding, TTL, correction policy or materialization lag. Count each cause separately. An aggregate parity percentage can hide a severe mismatch among new merchants or one high-risk segment. Report age distribution and misses alongside exact or tolerance-based value comparisons.
Test boundary cases
Include a late event, a duplicate delivery, a deleted entity, a null value and an online timeout. The audit should show which cases are expected divergence and which are defects. A model may have an allowed freshness window, but the test must not declare any amount of delay acceptable after observing the result.
Set a release gate
Run parity checks before promoting a model or changing a feature view. If the mismatch exceeds the predeclared tolerance, block promotion and attach a reproducible case with raw event IDs and both transform versions. Promotion gates should include feature health, because a sound model can fail when its input contract drifts.
Implementation
def parity_mismatches(online_values, replay_values):
mismatches = []
for decision_id, online_record in online_values.items():
replay_record = replay_values.get(decision_id)
if replay_record != online_record:
mismatches.append(decision_id)
return mismatchesPerformance and operating cost
Comparing D decision records costs O(D) expected time and O(M) output space for M mismatches when both sides are keyed. Replaying features may cost far more; use bounded samples and preserve enough lineage to reproduce each mismatch.
Common Mistakes
- Do not compare online now with an unrelated historical timestamp.
- Do not report parity without null and freshness comparisons.
- Do not waive a severe subgroup mismatch behind a strong pooled rate.
