Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Training-serving parity: compare feature values at one prediction clock

Last updated: 6 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A production feature must represent the same information and transformation that was available to the offline training example.

Define a feature view

Name source fields, types, null policy, event-time cutoff and transform version. A seven-day receipt count at 09:00 cannot include a transaction that posts at 09:05, even if the warehouse later backfills it. Compare offline and online values for the same entity and decision timestamp. Feature availability defines what the model was allowed to know.

Test both paths

Build a parity fixture with late arrivals, corrected amounts and absent history. Run it through offline and online transforms, then compare numeric values within justified tolerance. A shared function is helpful but does not prove parity if the paths read different source snapshots. Event-time policy must be shared too.

Name freshness separately

An online feature can match the offline formula yet be stale because its cache has not refreshed. Record source event time and serving refresh time, set a maximum age and choose a fallback for missing or stale values. A fallback of zero changes meaning when zero is a valid measured count; use an explicit missing indicator when required.

Monitor skew by cause

Log aggregate parity differences from sampled replayable requests without exposing sensitive records. Partition mismatches by feature version, source lag and entity type. A sudden mismatch after deployment may be a parser change, not model drift. Pause promotion or route to a safe fallback until the cause is understood.

Implementation

python
def compare_feature_snapshots(offline, online, tolerance=1e-6):
    if offline.keys() != online.keys():
        raise ValueError("feature names differ")
    mismatches = {}
    for feature_name in offline:
        left, right = offline[feature_name], online[feature_name]
        if left is None or right is None:
            different = left != right
        else:
            different = abs(left - right) > tolerance
        if different:
            mismatches[feature_name] = (left, right)
    return mismatches

Performance and operating cost

Comparing F scalar features costs O(F) per sampled request. The expensive work is reconstructing point-in-time offline values and retaining enough history to explain differences.

Common Mistakes

  • Do not compare values computed at different prediction times.
  • Do not mistake a stale cache for a changed formula.
  • Do not use zero as a universal missing fallback.

Read next

Continue the workflow: Vision augmentation: preserve labels and match serving transforms.

Continue the workflow: Feature parity: compare historical replay with online observations.

Continue the workflow: Feature contracts: admit only usable inference records.

Continue the workflow: Online feature freshness: use event and availability clocks.

Continue the workflow: Multi-stage inference: pin each stage and its contract.

Continue the workflow: Stream inference: event time, feature state and decision identity.

ai-data
mlops
Storage details