Training rows must retrieve the feature value that would have been visible when each prediction was made.
Historical feature retrieval: point-in-time joins without future leakage
Build entity rows first
Create one row per historical decision with entity ID, decision time, label horizon and an immutable row ID. Fetch features against those rows, not against the latest warehouse snapshot. A conventional as-of join chooses a value whose event time precedes the decision, but it must also respect availability time when ingestion lag or backfills exist. The feature contract defines both clocks.
Handle ties and corrections
Multiple values can share the same event time. Choose by a deterministic correction sequence or availability timestamp and record that rule. A later corrected value should not replace the original in a training row whose decision predates the correction. Set a lookback limit so a 14-month-old value is not silently used as current information.
Preserve missingness
Some entities have no prior feature. Return null with a missing indicator and declared default behavior instead of borrowing a later value. A missing rate that differs sharply between training and serving may indicate delayed materialization or a bad entity key. Parity checks should compare both values and missing markers for sampled decisions.
Use a two-row fixture
Merchant M has a count of 5 available at 08:45, a count of 8 whose event time is 09:00 but availability is 09:12, and decisions at 09:05 and 09:20. The first decision sees 5; the second can see 8. Add a correction available at 10:00 and verify it cannot rewrite the earlier rows.
Implementation
def point_in_time_value(feature_rows, entity_id, decision_time):
matches = [row for row in feature_rows
if row["entity_id"] == entity_id
and row["event_at"] <= decision_time
and row["available_at"] <= decision_time]
return max(matches, key=lambda row: (row["event_at"], row["available_at"]), default=None)Performance and operating cost
A naive join over R entity rows and F feature rows costs O(R times F). Partitioning by entity and indexing event and availability times reduces candidate scans; materializing historical training sets consumes storage but makes a run reproducible.
Common Mistakes
- Do not join historical labels to the newest feature row.
- Do not let a late correction alter an older prediction snapshot.
- Do not hide absent prior values with a future fill.
Read next
- Feature-store contract: entity keys, event time and availability
- Feature parity: compare historical replay with online observations
- Warehouse history: join facts to the dimension version valid at event time
- Prediction-time feature availability: reject future information before training
Continue the workflow: Table snapshots and atomic publication.
