Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Failover drills: reconcile decisions across regions

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A failover drill must prove one logical outcome per decision key even when clients retry across regions.

Give requests a stable identity

A client may time out after the primary writes a decision and retry against the standby. The retry should reuse the same idempotency key, scoped to caller and logical operation. A region-specific request ID is still useful for traces, but it cannot serve as the only deduplication key. Record the decision state, model digest, route and issuing region behind that stable key. Restricted logging keeps raw receipt data out of the reconciliation store while preserving a joinable identity.

State the recovery contract

Choose recovery time and tolerable feature-state loss before a drill. Define whether a failover can serve the last approved model or must match the newest approved digest, and what to do with lagged features. Keep a single operator decision record for switch, hold, rollback and return to primary. Region identity checks guard the start of serving; they do not prove that two regions will not issue conflicting responses during a split or a delayed DNS transition.

Reconcile, then resume normal routing

Collect all decision events for the drill window from both regions and group them by stable key. Identical repeated responses are harmless if consumers see one committed logical decision. Different routes or digests for one key require quarantine and review. Also compare gateway counts with decision-log counts to find lost events. Preserve a record of requests with no durable response; do not invent an outcome from a log gap. Outcome joins later depend on these IDs remaining stable.

Measure what the drill changed

Report switch time, rejected or manually reviewed receipts, feature lag, duplicate retries, divergent decisions, log gaps and return-to-primary time. Test a false-positive health alarm as well as an actual regional loss, because unnecessary switches create their own risk. The project demonstrates an apparently successful cutover that still fails reconciliation because one receipt was scored differently on both sides. The drill is complete only when the discrepancy has an owner and a resolution.

Implementation

python
def reconcile_decisions(events):
    grouped = {}
    for event in events:
        grouped.setdefault(event["decision_key"], set()).add(
            (event["route"], event["model_digest"]))
    divergent = sorted(key for key, outcomes in grouped.items()
                       if len(outcomes) > 1)
    return {"state": "hold" if divergent else "consistent",
            "divergent_keys": divergent}

events = [{"decision_key": "receipt-47", "route": "review",
           "model_digest": "model-47"},
          {"decision_key": "receipt-47", "route": "review",
           "model_digest": "model-47"}]
assert reconcile_decisions(events)["state"] == "consistent"
assert reconcile_decisions(events + [{**events[0], "route": "clear"}])[
    "divergent_keys"] == ["receipt-47"]

Performance and operating cost

For n events, grouping uses O(n) expected time and O(n) memory; sorting only divergent keys adds O(d log d) time for d discrepancies. Cross-region logs may arrive late, so the drill must wait for a defined watermark before declaring no divergence. The example checks response identity, not whether the underlying business action was applied twice.

Common Mistakes

  • Generating a new idempotency key on every retry.
  • Calling the drill successful from gateway uptime alone.
  • Dropping logs from the old region before reconciliation.
  • Assuming identical API responses prove downstream side effects were deduplicated.

Read next

Continue the workflow: Inference retries: bound repeated work and preserve one decision.

ai-data
mlops
Storage details