A failover drill must prove one logical outcome per decision key even when clients retry across regions.
Failover drills: reconcile decisions across regions
Give requests a stable identity
A client may time out after the primary writes a decision and retry against the standby. The retry should reuse the same idempotency key, scoped to caller and logical operation. A region-specific request ID is still useful for traces, but it cannot serve as the only deduplication key. Record the decision state, model digest, route and issuing region behind that stable key. Restricted logging keeps raw receipt data out of the reconciliation store while preserving a joinable identity.
State the recovery contract
Choose recovery time and tolerable feature-state loss before a drill. Define whether a failover can serve the last approved model or must match the newest approved digest, and what to do with lagged features. Keep a single operator decision record for switch, hold, rollback and return to primary. Region identity checks guard the start of serving; they do not prove that two regions will not issue conflicting responses during a split or a delayed DNS transition.
Reconcile, then resume normal routing
Collect all decision events for the drill window from both regions and group them by stable key. Identical repeated responses are harmless if consumers see one committed logical decision. Different routes or digests for one key require quarantine and review. Also compare gateway counts with decision-log counts to find lost events. Preserve a record of requests with no durable response; do not invent an outcome from a log gap. Outcome joins later depend on these IDs remaining stable.
Measure what the drill changed
Report switch time, rejected or manually reviewed receipts, feature lag, duplicate retries, divergent decisions, log gaps and return-to-primary time. Test a false-positive health alarm as well as an actual regional loss, because unnecessary switches create their own risk. The project demonstrates an apparently successful cutover that still fails reconciliation because one receipt was scored differently on both sides. The drill is complete only when the discrepancy has an owner and a resolution.
Implementation
def reconcile_decisions(events):
grouped = {}
for event in events:
grouped.setdefault(event["decision_key"], set()).add(
(event["route"], event["model_digest"]))
divergent = sorted(key for key, outcomes in grouped.items()
if len(outcomes) > 1)
return {"state": "hold" if divergent else "consistent",
"divergent_keys": divergent}
events = [{"decision_key": "receipt-47", "route": "review",
"model_digest": "model-47"},
{"decision_key": "receipt-47", "route": "review",
"model_digest": "model-47"}]
assert reconcile_decisions(events)["state"] == "consistent"
assert reconcile_decisions(events + [{**events[0], "route": "clear"}])[
"divergent_keys"] == ["receipt-47"]
Performance and operating cost
For n events, grouping uses O(n) expected time and O(n) memory; sorting only divergent keys adds O(d log d) time for d discrepancies. Cross-region logs may arrive late, so the drill must wait for a defined watermark before declaring no divergence. The example checks response identity, not whether the underlying business action was applied twice.
Common Mistakes
- Generating a new idempotency key on every retry.
- Calling the drill successful from gateway uptime alone.
- Dropping logs from the old region before reconciliation.
- Assuming identical API responses prove downstream side effects were deduplicated.
Read next
- Region failover for inference: match model and feature state
- Project: fail over receipt inference and reconcile every decision
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Prediction-outcome joins: evaluate only mature, matched decisions
- Model incidents: build a release and evidence timeline before rollback
Continue the workflow: Inference retries: bound repeated work and preserve one decision.
