Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Dependency readiness and source freshness

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A source is ready for a pipeline interval when its declared completeness condition is met; freshness says how far the observed data trails the expected event or ingest clock.

Do not confuse present with complete

A day-47 object may exist while its uploader has sent only 60% of the batch. An empty interval may be valid if the producer explicitly marks it complete. Prefer a producer manifest containing interval, expected files or count, checksum and completion timestamp. A polling task should wait for the manifest and verify it, not just test for one file path.

Choose the clock

Ingest freshness measures when the platform last received data; event freshness measures when the newest business event occurred. A source with frequent retries can appear ingest-fresh while its latest business event is days old. Record both timestamps and the expected cadence. Late-event policy states whether an interval may receive corrections after initial completion.

Set a useful threshold

If a dispatch dashboard is promised by 08:30, a source arriving at 08:24 leaves six minutes for transformation and validation. A readiness threshold of 08:29 is too late. Work backward from consumer deadline, p95 processing time, and an operational margin. Alert on a missed readiness boundary before the dashboard itself is late.

Express blocked states

A pipeline waiting on a late source is blocked, not failed by its own computation. Record the missing interval, producer owner, last observed event time and deadline. Retry checks at a bounded rate. When the source completes, run against its pinned version; if it misses the cutoff, follow a documented stale-data or no-publish policy rather than silently reusing yesterday's data.

Keep readiness separate from quality

A completed batch of 2,700 records may still contain impossible amounts or duplicate keys. Completeness proves that extraction finished; quality gates decide whether the records are publishable. Both statuses must appear in the release manifest.

Implementation

python
source_manifest = {"interval": 47, "expected_files": 3,
                   "observed_files": 3, "complete": True,
                   "latest_event_minute": 474}

def source_status(manifest, required_interval, newest_allowed_minute):
    if manifest["interval"] != required_interval:
        return "wrong_interval"
    if not manifest["complete"] or manifest["observed_files"] != manifest["expected_files"]:
        return "incomplete"
    if manifest["latest_event_minute"] < newest_allowed_minute:
        return "stale"
    return "ready"

assert source_status(source_manifest, 47, 470) == "ready"
assert source_status(source_manifest, 47, 480) == "stale"

Performance and operating cost

A manifest check is O(1) for fixed summary fields, plus O(F) to verify F file checksums when required. Polling too often raises storage or API cost without making an unavailable source arrive sooner. Set cadence from the deadline and source behavior.

Common Mistakes

  • Do not use file existence as proof of complete delivery.
  • Do not measure only ingest freshness when business events can stall.
  • Do not publish stale data without labeling the cutoff and consumer policy.

Read next

ai-data
data-engineering
Storage details