Design a case-review service with readers in regions 29 and 47 and one authoritative write region. Public hashed assets can be served near users, but a private case response must remain bound to authenticated tenant rights and never be shared by URL alone. A save returns a revision token. If a nearby replica is behind that revision, the next read waits within a bound, routes to the writer, or reports a pending state; it must not silently replace the editor with stale data. A recovery drill promotes a second region only after old writers are fenced and possible unreplicated writes are recorded. Browser retries reuse idempotency keys, while queued email and export work resume from durable intents after reconciliation. Map where scans, derivatives, caches, logs, backups, and support exports reside. Tenant placement policy applies to the recovery destination as well as normal requests. The project is an architecture and failure test; service guarantees must be checked against the selected database and network provider.
Project: two-region case service recovery
Build contract
- Warm a cache as one tenant and verify a second tenant cannot read the response.
- Save revision 47 while a replica remains at 46 and exercise the freshness fallback.
- Promote generation 7 while rejecting generation 6 writers and reconciling retries.
- Inventory private data, derived files, telemetry, backups, and allowed recovery regions.
Implementation checkpoint
function readFromReplica(appliedRevision, requiredRevision, regionAllowed) {
return regionAllowed && appliedRevision >= requiredRevision;
}
console.log(readFromReplica(46, 47, true));
// Output: falseCost and boundaries
Public caching can reduce asset latency; private freshness checks may add a cross-region hop or bounded wait. A passive second region consumes storage and drill effort; an active one adds compute and harder coordination. Asynchronous replication creates a possible loss window at failover, so record lag and accepted write generations. A regional placement rule can lower cache efficiency and restrict recovery choices. Measure user-visible stale reads, fallback-to-writer share, promotion time, replay count, and unclassified private data flows.
Failure drill
Try a private URL warmed under another tenant, then shift edge traffic and repeat. Stop replication at revision 46 after the writer acknowledges 47; verify the client does not show 46 as final. Leave an old batch worker alive after promotion and assert its writes fail. Replay a timed-out save using the original key and inspect the outbox before sending email again. Attempt recovery into a region disallowed for the tenant and inspect logs and temporary exports for stray private content.
Acceptance checks
- Private responses never cross permission scope.
- Post-save reads meet a named freshness contract.
- One writer generation owns mutations after failover.
- Placement inventory covers derivatives and recovery copies.
Common Mistakes
- Treating nearest edge as the database authority.
- Using a traffic switch as the only writer fence.
- Assuming one regional database setting covers logs and backups.
Related lessons
Regional Routing, Static and Private Cache Boundaries; Replica Lag, Read-Your-Write, and Version Cursors; Regional Failover, Fencing, and Replay; Regional Data Location and Operational Evidence.
