A source is ready for a pipeline interval when its declared completeness condition is met; freshness says how far the observed data trails the expected event or ingest clock.
Dependency readiness and source freshness
Do not confuse present with complete
A day-47 object may exist while its uploader has sent only 60% of the batch. An empty interval may be valid if the producer explicitly marks it complete. Prefer a producer manifest containing interval, expected files or count, checksum and completion timestamp. A polling task should wait for the manifest and verify it, not just test for one file path.
Choose the clock
Ingest freshness measures when the platform last received data; event freshness measures when the newest business event occurred. A source with frequent retries can appear ingest-fresh while its latest business event is days old. Record both timestamps and the expected cadence. Late-event policy states whether an interval may receive corrections after initial completion.
Set a useful threshold
If a dispatch dashboard is promised by 08:30, a source arriving at 08:24 leaves six minutes for transformation and validation. A readiness threshold of 08:29 is too late. Work backward from consumer deadline, p95 processing time, and an operational margin. Alert on a missed readiness boundary before the dashboard itself is late.
Express blocked states
A pipeline waiting on a late source is blocked, not failed by its own computation. Record the missing interval, producer owner, last observed event time and deadline. Retry checks at a bounded rate. When the source completes, run against its pinned version; if it misses the cutoff, follow a documented stale-data or no-publish policy rather than silently reusing yesterday's data.
Keep readiness separate from quality
A completed batch of 2,700 records may still contain impossible amounts or duplicate keys. Completeness proves that extraction finished; quality gates decide whether the records are publishable. Both statuses must appear in the release manifest.
Implementation
source_manifest = {"interval": 47, "expected_files": 3,
"observed_files": 3, "complete": True,
"latest_event_minute": 474}
def source_status(manifest, required_interval, newest_allowed_minute):
if manifest["interval"] != required_interval:
return "wrong_interval"
if not manifest["complete"] or manifest["observed_files"] != manifest["expected_files"]:
return "incomplete"
if manifest["latest_event_minute"] < newest_allowed_minute:
return "stale"
return "ready"
assert source_status(source_manifest, 47, 470) == "ready"
assert source_status(source_manifest, 47, 480) == "stale"Performance and operating cost
A manifest check is O(1) for fixed summary fields, plus O(F) to verify F file checksums when required. Polling too often raises storage or API cost without making an unavailable source arrive sooner. Set cadence from the deadline and source behavior.
Common Mistakes
- Do not use file existence as proof of complete delivery.
- Do not measure only ingest freshness when business events can stall.
- Do not publish stale data without labeling the cutoff and consumer policy.
Read next
- DAG intervals and idempotent task outputs
- Data quality gates: quarantine bad rows and reconcile complete batches
- Pipeline SLOs, freshness and error budgets
- Replay manifests and audit trails
- Project: operate a daily billing pipeline
- Event time and late arrivals: close windows with an explicit correction policy
