Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Refresh lag and cost budgets

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A refresh budget joins an observed staleness limit with compute and failure limits, so a view is neither perpetually stale nor needlessly rebuilt.

Define the consumer deadline

A dashboard may tolerate a 12-minute-old store-day total, while fraud operations need a fresher feed. Record the maximum age at the point of query, the consequence of stale output and whether a dated last-good value is acceptable. The allowed lag must include ingestion, upstream processing, view refresh and index publication; assigning the full budget to the final job hides upstream delay.

Measure stage age separately

Track newest accepted source position, candidate refresh start, last committed view generation and current reader pointer. A view can be freshly rebuilt from old source data. Emit both source age and view age, and label the interval when a correction window is still open. Pipeline SLOs should count the consumer-visible failure, not only successful scheduler invocations.

Choose a policy from cost curves

Shorter intervals increase fixed startup work, metadata traffic and contention; longer intervals increase staleness and the size of each change set. Measure bytes scanned, changed groups, write volume and wall time per refresh. Use these observations to choose a target and capacity margin. When a backlog appears, prioritize the newest useful generation and cancel superseded work only if publication semantics permit.

Degrade explicitly under stress

If the next refresh would exceed its deadline, keep a dated last-good snapshot and raise a freshness breach. Do not replace a correct but stale answer with an incomplete candidate. Define whether a full rebuild is allowed during peak traffic and reserve a recovery budget. Admission control can protect interactive reads while a large repair runs.

Exercise a burst

Generate 47 change batches while a refresh worker is paused, then resume it. Observe queue depth, actual age, scanned bytes and first good consumer query. Inject one failed candidate followed by a successful generation. Verify that a green job status without a moved serving pointer does not clear the breach. Record the cost per recovered minute of freshness.

Implementation

python
from datetime import datetime, timezone

def visible_age_seconds(now, source_committed_at, view_built_at):
    return {"source_age": (now - source_committed_at).total_seconds(),
            "view_age": (now - view_built_at).total_seconds()}

clock = datetime(2026, 10, 6, 12, 0, tzinfo=timezone.utc)
ages = visible_age_seconds(clock,
    datetime(2026, 10, 6, 11, 47, tzinfo=timezone.utc),
    datetime(2026, 10, 6, 11, 58, tzinfo=timezone.utc))
assert ages == {"source_age": 780.0, "view_age": 120.0}
assert ages["source_age"] > 720

Performance and operating cost

The age check is O(1) time and space. Refresh cost depends on bytes scanned, changed groups and view write amplification, while telemetry cost grows with retained samples and workload dimensions. A fixed worker reservation costs capacity even when idle; the alternative is a documented recovery delay during bursts. Alert on observed age rather than an ideal schedule.

Common Mistakes

  • Do not infer fresh source data from a recent view build timestamp.
  • Do not clear a freshness incident before the reader pointer moves.
  • Do not shorten a target lag without checking whether refresh runtime can meet it.

Read next

ai-data
data-engineering
Storage details