Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Pipeline RPO, RTO and replication lag

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A recovery point objective bounds acceptable data loss; a recovery time objective bounds how long a data product may remain unavailable after a failure.

Define the product boundary

A replicated object store alone does not make a complete analytics service recoverable. Include source offsets, checkpoint state, table metadata, access policy, schema versions and serving configuration in the recovery unit. If a replica has files but not the snapshot pointer that names them, the reader cannot safely publish the newest data. The release manifest should identify every required component.

Measure lag at the last consistent point

A monitor reporting five minutes of object replication lag does not prove five-minute RPO if table metadata or checkpoints trail by 42 minutes. Track the newest complete generation that can be opened and validated in the recovery region. RPO is the age of that generation at failure time. Record source high-water marks and a digest for each release.

Include activation in RTO

Recovery time starts with detection and includes operator decision, credential availability, restore, data validation, route switch and the first successful consumer query. A runbook that times only instance startup understates the outage. Define what degraded mode is acceptable, such as serving a clearly dated last-good snapshot while new ingestion catches up.

Avoid split-brain writes

If both regions resume publishing after a network partition, two independent snapshot histories can emerge. Use a single writer lease with a monotonically increasing epoch or another fencing mechanism. Fenced promotion prevents the old writer from committing after the recovery writer takes control.

Prove the target in a drill

Disable the primary in a controlled exercise, restore from a selected generation and run real consumer queries. Measure last consistent input position, actual missing events, elapsed recovery time and post-restore backlog. A backup is evidence of stored bytes; a successful query and reconciliation are evidence of a usable data product.

Implementation

python
from datetime import datetime, timezone

failure_at = datetime(2026, 10, 6, 12, 0, tzinfo=timezone.utc)
last_complete = datetime(2026, 10, 6, 11, 37, tzinfo=timezone.utc)
first_query = datetime(2026, 10, 6, 12, 41, tzinfo=timezone.utc)

def recovery_minutes(failure, recoverable, query_ready):
    if not recoverable <= failure <= query_ready:
        raise ValueError("recovery times are out of order")
    return ((failure - recoverable).total_seconds() / 60,
            (query_ready - failure).total_seconds() / 60)

assert recovery_minutes(failure_at, last_complete, first_query) == (23, 41)

Performance and operating cost

The arithmetic is O(1), but the recovery design pays for replicated bytes, retained checkpoints, standby compute, drills and catch-up processing. A smaller RPO generally requires faster complete-generation replication; a smaller RTO requires more prepared capacity and fewer manual activation steps. Measure both against a working consumer query.

Common Mistakes

  • Do not use file replication lag as the only measure of recoverable data age.
  • Do not time RTO from instance boot instead of failure detection.
  • Do not promote a second writer without fencing the first.

Read next

Continue the workflow: Regional data boundaries and egress gates.

ai-data
data-engineering
Storage details