A production feature must represent the same information and transformation that was available to the offline training example.
Training-serving parity: compare feature values at one prediction clock
Define a feature view
Name source fields, types, null policy, event-time cutoff and transform version. A seven-day receipt count at 09:00 cannot include a transaction that posts at 09:05, even if the warehouse later backfills it. Compare offline and online values for the same entity and decision timestamp. Feature availability defines what the model was allowed to know.
Test both paths
Build a parity fixture with late arrivals, corrected amounts and absent history. Run it through offline and online transforms, then compare numeric values within justified tolerance. A shared function is helpful but does not prove parity if the paths read different source snapshots. Event-time policy must be shared too.
Name freshness separately
An online feature can match the offline formula yet be stale because its cache has not refreshed. Record source event time and serving refresh time, set a maximum age and choose a fallback for missing or stale values. A fallback of zero changes meaning when zero is a valid measured count; use an explicit missing indicator when required.
Monitor skew by cause
Log aggregate parity differences from sampled replayable requests without exposing sensitive records. Partition mismatches by feature version, source lag and entity type. A sudden mismatch after deployment may be a parser change, not model drift. Pause promotion or route to a safe fallback until the cause is understood.
Implementation
def compare_feature_snapshots(offline, online, tolerance=1e-6):
if offline.keys() != online.keys():
raise ValueError("feature names differ")
mismatches = {}
for feature_name in offline:
left, right = offline[feature_name], online[feature_name]
if left is None or right is None:
different = left != right
else:
different = abs(left - right) > tolerance
if different:
mismatches[feature_name] = (left, right)
return mismatchesPerformance and operating cost
Comparing F scalar features costs O(F) per sampled request. The expensive work is reconstructing point-in-time offline values and retaining enough history to explain differences.
Common Mistakes
- Do not compare values computed at different prediction times.
- Do not mistake a stale cache for a changed formula.
- Do not use zero as a universal missing fallback.
Read next
- Prediction-time feature availability: reject future information before training
- Event time and late arrivals: close windows with an explicit correction policy
- Model monitoring: separate input drift, data faults and delayed outcomes
- Model promotion: require evidence before changing the serving pointer
Continue the workflow: Vision augmentation: preserve labels and match serving transforms.
Continue the workflow: Feature parity: compare historical replay with online observations.
Continue the workflow: Feature contracts: admit only usable inference records.
Continue the workflow: Online feature freshness: use event and availability clocks.
Continue the workflow: Multi-stage inference: pin each stage and its contract.
Continue the workflow: Stream inference: event time, feature state and decision identity.
