Recover a customer-day mart in a secondary region while preserving source position, writer ownership and the serving view’s generation.
Project: recover a regional analytics mart
Set measurable recovery targets
The product publishes a customer-day count and net cents every 24 minutes. Set an RPO of 36 minutes and an RTO of 68 minutes for the exercise, then define which consumer query proves recovery. The manifest must name source positions, table snapshot, schema versions, index generation and access-policy revision. The RPO/RTO contract applies to that complete set.
Prepare the replica
Copy data files and table metadata to the secondary region, verify checksums, and retain the source log past the worst expected outage and catch-up duration. Build an index generation from each complete table snapshot. Do not publish a replica merely because files arrived; the metadata pointer and access policy must agree with them.
Force a clean promotion
Stop primary writes at a chosen source position, simulate its loss and grant the secondary a higher writer epoch. Restore from the newest complete generation, replay missing source records and atomically publish the next table and index pointers. Fencing must reject a delayed primary commit after promotion.
Test the consumer path
Run account-day queries before failure, during last-good serving and after recovery. Every answer should name one index generation and its source cutoff. Compare aggregate cents and unique account-day keys with the source manifest; verify access roles still filter rows and mask fields correctly in the second region.
Record the incident evidence
Deliver timestamps for failure detection, last recoverable generation and first successful query; calculate actual RPO and RTO. Include rejected old-writer attempt, replayed source interval, orphan candidates, index switch and cost of the drill. Reconcile the returning primary before any failback. A green infrastructure dashboard is not sufficient evidence of correct answers.
Implementation
from datetime import datetime, timezone
failure = datetime(2026, 10, 6, 12, 0, tzinfo=timezone.utc)
complete = datetime(2026, 10, 6, 11, 35, tzinfo=timezone.utc)
query_ready = datetime(2026, 10, 6, 12, 54, tzinfo=timezone.utc)
target_rpo_minutes = 36
target_rto_minutes = 68
actual_rpo = (failure - complete).total_seconds() / 60
actual_rto = (query_ready - failure).total_seconds() / 60
assert actual_rpo == 25 and actual_rto == 54
assert actual_rpo <= target_rpo_minutes
assert actual_rto <= target_rto_minutesPerformance and operating cost
The objective calculation is O(1), while the release costs include cross-region transfer, retained snapshots, replay compute, standby serving capacity and periodic drills. A full index rebuild is O(N) in customer-day rows and can dominate the measured RTO, so include it in the timed exercise rather than declaring recovery after the table alone opens.
Common Mistakes
- Do not report RPO from the newest copied file instead of a complete generation.
- Do not let the old region resume writes after secondary promotion.
- Do not count a restored table as recovered while the consumer index or access policy is missing.
Read next
- Pipeline RPO, RTO and replication lag
- Failover writer fencing and source positions
- Serving indexes and freshness contracts
- Replay manifests and audit trails
- Row policy and column mask tests
Continue the workflow: Project: prove erasure after a regional restore.
Continue the workflow: Project: protect analytics reads during a tenant burst.
