Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: recover a regional analytics mart

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Recover a customer-day mart in a secondary region while preserving source position, writer ownership and the serving view’s generation.

Set measurable recovery targets

The product publishes a customer-day count and net cents every 24 minutes. Set an RPO of 36 minutes and an RTO of 68 minutes for the exercise, then define which consumer query proves recovery. The manifest must name source positions, table snapshot, schema versions, index generation and access-policy revision. The RPO/RTO contract applies to that complete set.

Prepare the replica

Copy data files and table metadata to the secondary region, verify checksums, and retain the source log past the worst expected outage and catch-up duration. Build an index generation from each complete table snapshot. Do not publish a replica merely because files arrived; the metadata pointer and access policy must agree with them.

Force a clean promotion

Stop primary writes at a chosen source position, simulate its loss and grant the secondary a higher writer epoch. Restore from the newest complete generation, replay missing source records and atomically publish the next table and index pointers. Fencing must reject a delayed primary commit after promotion.

Test the consumer path

Run account-day queries before failure, during last-good serving and after recovery. Every answer should name one index generation and its source cutoff. Compare aggregate cents and unique account-day keys with the source manifest; verify access roles still filter rows and mask fields correctly in the second region.

Record the incident evidence

Deliver timestamps for failure detection, last recoverable generation and first successful query; calculate actual RPO and RTO. Include rejected old-writer attempt, replayed source interval, orphan candidates, index switch and cost of the drill. Reconcile the returning primary before any failback. A green infrastructure dashboard is not sufficient evidence of correct answers.

Implementation

python
from datetime import datetime, timezone

failure = datetime(2026, 10, 6, 12, 0, tzinfo=timezone.utc)
complete = datetime(2026, 10, 6, 11, 35, tzinfo=timezone.utc)
query_ready = datetime(2026, 10, 6, 12, 54, tzinfo=timezone.utc)
target_rpo_minutes = 36
target_rto_minutes = 68

actual_rpo = (failure - complete).total_seconds() / 60
actual_rto = (query_ready - failure).total_seconds() / 60
assert actual_rpo == 25 and actual_rto == 54
assert actual_rpo <= target_rpo_minutes
assert actual_rto <= target_rto_minutes

Performance and operating cost

The objective calculation is O(1), while the release costs include cross-region transfer, retained snapshots, replay compute, standby serving capacity and periodic drills. A full index rebuild is O(N) in customer-day rows and can dominate the measured RTO, so include it in the timed exercise rather than declaring recovery after the table alone opens.

Common Mistakes

  • Do not report RPO from the newest copied file instead of a complete generation.
  • Do not let the old region resume writes after secondary promotion.
  • Do not count a restored table as recovered while the consumer index or access policy is missing.

Read next

Continue the workflow: Project: prove erasure after a regional restore.

Continue the workflow: Project: protect analytics reads during a tenant burst.

ai-data
data-engineering
Storage details