Build a small offline and online feature path for merchant receipt triage, then prove historical availability and serving parity.
Project: build a point-in-time receipt-triage feature service
Write contracts and fixture
Define a merchant key, a 47-hour rejected-receipt count and a last-submission-age feature. Give raw events event time, ingestion time, immutable event ID and correction version. Add a late event, a retry, a renamed merchant, a deletion and a new merchant with no history. The feature contract must state units, missing policy, owner and prediction cutoff.
Build historical retrieval
Create decision rows at two times for each merchant. Retrieve only values whose event and availability times meet each decision cutoff, with deterministic tie handling. Show one row where a late event would leak if only event time were checked. Compare your output with a deliberately naive latest-value join and explain the difference for a specific decision ID.
Serve and validate
Materialize a versioned online snapshot. Return value, age, version and missing or stale status for each request. Add a TTL, deletion gate and rollback path. Sample request traces and replay the corresponding historical decisions; report exact parity, missing-state parity and mismatch reasons. Use a small fixture, but keep the same cutoff logic in the replay and serving code.
Deliver a release packet
Submit schema, transform, fixture, decision rows, backfill manifest, parity report and one deletion-race test. State the expected cost of each retrieval mode and the operational response to an online timeout. A reviewer should be able to reconstruct one training row and one online response from raw events and version IDs.
Implementation
def latest_visible_count(records, merchant_id, decision_minute):
eligible = [record for record in records
if record["merchant_id"] == merchant_id
and record["event_minute"] <= decision_minute
and record["available_minute"] <= decision_minute]
latest = max(eligible, key=lambda record: (record["event_minute"], record["available_minute"]), default=None)
return None if latest is None else latest["count"]Performance and operating cost
A fixture scan costs O(F) time and O(F) temporary space for F feature records per decision. A production historical join needs indexed partitions or a batch retrieval plan; online lookup needs bounded latency and versioned publication.
Common Mistakes
- Do not train from a feature value that arrived after the decision.
- Do not make a deleted merchant reappear during rollback.
- Do not mark stale online data as fresh because the key exists.
