A valid feature value can still be too old for a decision; freshness needs a task-specific limit and explicit missing state.
Online feature freshness: use event and availability clocks
Name both clocks
A merchant activity feature has an event time describing the activity and an availability time when the online store received it. A value can have a recent event time but arrive late, or an old event time that was just republished. Record both. For receipt-risk scoring, define a maximum age from event time and a publication lag limit from availability time. Training-serving parity compares values under a prediction clock; freshness adds a runtime rule for whether the value is usable now.
Separate TTL from safety
A store TTL may remove records after a configured duration. It does not prove values remaining in the store are fresh enough for every model. A daily merchant category can tolerate more age than a minute-level fraud signal. Enforce the model-specific freshness budget at retrieval, not only in storage configuration. A missing value and an expired value should have distinguishable reason codes, even if both route to the same approved fallback. Admission gates define that fallback boundary.
Choose behavior before an outage
When a feature is stale, the service may abstain, use a documented older model that does not require it or send the receipt for review. Do not silently fill the feature with the last cached value and report a normal score. Keep fallback volume and source lag visible by feature view and region. Avoid entity IDs as metric labels; investigate individual records through restricted traces. Serving fallback covers capacity failures, while this policy covers information age.
Test late and out-of-order writes
Feed an older event after a newer event, a future timestamp, a record exactly at the freshness boundary and an online-store outage. The latest arrival need not be the latest event. Check that a late correction updates historical evidence without overwriting a newer online value improperly. The freshness project tests both serving decisions and the monitoring signals that reveal a stuck materialization job.
Implementation
from datetime import timedelta
def feature_freshness(record, decision_at, max_age_minutes=47):
if record is None:
return {"state": "unavailable", "reason": "missing"}
event_at = record["event_at"]
available_at = record["available_at"]
if event_at > decision_at or available_at > decision_at:
return {"state": "unavailable", "reason": "future-clock"}
if decision_at - event_at > timedelta(minutes=max_age_minutes):
return {"state": "unavailable", "reason": "stale-event"}
return {"state": "usable", "value": record["value"]}
from datetime import datetime, timezone
now = datetime(2026, 10, 6, tzinfo=timezone.utc)
recent = {"event_at": now - timedelta(minutes=23),
"available_at": now - timedelta(minutes=4), "value": 82}
assert feature_freshness(recent, now)["state"] == "usable"
assert feature_freshness({**recent, "event_at": now - timedelta(minutes=82)},
now)["reason"] == "stale-event"
Performance and operating cost
One feature check is O(1) time and space. Fetching many feature views adds network latency and monitoring dimensions. A shorter TTL can reduce stale reads but increase misses and recomputation; choose it from decision needs and source update behavior, not as a substitute for a request-time freshness check.
Common Mistakes
- Treating a stored value as fresh merely because TTL has not deleted it.
- Using ingestion time in place of event time.
- Silently scoring with a stale cache value.
- Allowing an out-of-order older event to replace a newer online value.
Read next
- Feature materialization lag: detect stuck updates and repair safely
- Project: keep receipt scoring safe during feature publication lag
- Training-serving parity: compare feature values at one prediction clock
- Feature contracts: admit only usable inference records
- Serving overload: bound queues and choose a fallback before time runs out
Continue the workflow: Region failover for inference: match model and feature state.
Continue the workflow: Prediction caches: key by every input that changes the decision.
