DNS caches, persistent HTTP connections, proxy pools, mobile clients, and queued jobs do not all change endpoint at once. A drain contract defines which requests can finish on the old region, which receive a retryable refusal, how long an old connection may persist, and how a repeated operation is identified. The old region must keep enforcing the new write epoch throughout convergence; draining without fencing leaves a narrow but real second-writer window.
Cutover drains: handle old connections and ambiguous client retries
Operational decision
A parcel-label API routes new sessions to a standby while old mobile clients keep established connections. Read-only label lookups can complete on the old endpoint if their staleness bound is acceptable. Label purchases require the current write epoch; a stale connection receives a retryable response with no side effect. Clients reuse the original purchase idempotency key when retrying against the standby. Instrument requests by region, operation, client connection age, epoch, and outcome, then wait for the old-region in-flight count to reach the agreed drain threshold. Do not kill every old connection immediately: an abrupt close during a response can make clients uncertain whether a purchase committed. Also do not wait forever for one idle client; enforce a maximum age and a safe retry path. Test a request that commits just before connection loss and a request rejected just after epoch change. Both must have a discoverable final state. If a background worker uses a private endpoint that bypasses public routing, include it in the drain inventory and keep it fenced until its owner confirms the new destination.
Parcel cutover matrix
Old read: serve only within staleness contract
Old write with stale epoch: reject without side effect
New write: accept with stable idempotency key
Ambiguous response: query by request ID, then retry
Private worker: switch endpoint and verify epoch
Drain complete: old in-flight below thresholdCost and verification
Connection age and in-flight counts require per-region metrics and enough cardinality control to avoid exploding telemetry cost. A long drain holds capacity and certificates in two places; a short drain increases retries and ambiguous outcomes. Measure successful customer operations, duplicate purchase attempts, old-region rejected writes, connection age, and time until the last old write-capable path is closed.
Common Mistakes
- Do not assume a routing-control change closes existing connections.
- Do not generate a new idempotency key on every retry.
- Do not forget private workers that bypass the public hostname.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- DNS cutovers: budget for resolver caches and mixed destinations
- Idempotency keys: reconcile an accepted write before repeating it
- Circuit breaker recovery: probe capacity without reopening a flood
- Failover fencing: prevent two writable database leaders
