Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Pipeline SLOs, freshness and error budgets

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A pipeline service-level objective states the measurable share of intervals whose usable data arrives before a consumer deadline with the required quality.

Define a good interval

For the finance daily table, define success as a committed day partition by 08:30 local time, all required source manifests complete, accepted plus quarantined count equal to extracted count, and no critical quality violation. A scheduler's green task status is not enough. The user-facing unit is a usable interval, so measure it from the consumer's query path.

Calculate the budget

An objective of 99% good daily intervals over a 100-day rolling window permits one bad interval. A target of 99.9% on only 30 daily intervals is hard to interpret because one failure moves the observed rate by more than three percentage points. Select a window and unit with enough opportunities, or use a weekly or per-run target with clear meaning.

Separate symptoms

Track source lag, processing duration, quality rejection rate and publication delay alongside the final good-interval indicator. A late source and a slow transform require different owners. Readiness tells whether the source delivered; failure classification tells whether retrying helps.

Alert for action

Page a human when the consumer deadline is at risk and the team can act, not on every transient task retry. A warning can trigger at 07:50 when the source is missing; an incident can trigger at 08:30 when no valid snapshot is published. Include affected dataset, interval, run ID, last good snapshot and runbook path in the alert.

Use the budget in release decisions

If recent intervals exhausted the failure budget, postpone a risky schema change and repair the unstable pipeline first. The budget is not a reason to hide bad intervals or lower quality gates. A data product with a strict deletion or financial correctness requirement may need a separate zero-tolerance guardrail even when its freshness objective permits rare lateness.

Implementation

python
daily_intervals = [
    {"day": 47, "ready": True, "quality_ok": True, "published_minute": 506},
    {"day": 48, "ready": True, "quality_ok": True, "published_minute": 518},
    {"day": 49, "ready": False, "quality_ok": False, "published_minute": None},
]

def good_interval_rate(records, deadline_minute=510):
    if not records:
        raise ValueError("no intervals")
    good = sum(record["ready"] and record["quality_ok"]
               and record["published_minute"] is not None
               and record["published_minute"] <= deadline_minute
               for record in records)
    return good / len(records)

assert good_interval_rate(daily_intervals) == 1 / 3

Performance and operating cost

Aggregating N intervals is O(N) time and O(1) extra space. Monitoring cost comes from trustworthy source and publication timestamps, retention of interval outcomes, and alerts with enough context to act. A small denominator can make percentage targets misleading.

Common Mistakes

  • Do not measure only task success when the published table is unusable.
  • Do not set an SLO without a consumer deadline and quality condition.
  • Do not treat a freshness allowance as permission to publish wrong financial data.

Read next

Continue the workflow: Stream backpressure and catch-up budget.

Continue the workflow: Serving indexes and freshness contracts.

Continue the workflow: Telemetry correlation and cardinality budgets.

Continue the workflow: Project: gate a payment mart on quality evidence.

Continue the workflow: Refresh lag and cost budgets.

ai-data
data-engineering
Storage details