Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: fail over receipt inference and reconcile every decision

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Run an inference switch with an artifact mismatch, lagged features and a cross-region retry, then publish the recovery evidence.

Prepare both decision paths

Deploy the approved receipt artifact and API policy to two regions. Record their resolved digests, feature watermarks, contract versions and log sinks. Set a per-feature age budget and a recovery-time objective. Keep the standby out of the live route until its canary verifies a real decision response with the expected identity. The readiness gate must distinguish “endpoint responds” from “safe to score.”

Inject failure in stages

First interrupt the primary after it logs a decision but before the client receives the response. Retry the same logical receipt against the standby with the same idempotency key. Next stage an old artifact on standby; expect a hold. Finally restore the artifact but delay its feature stream beyond the permitted age; expect manual review or rejection. Do not weaken the gate to make the drill pass. Record every route switch and its operator.

Reconcile the overlap window

Export redacted decisions and gateway requests from both regions for a fixed window that includes late log arrival. Group by stable key, compare route and model digest, and flag missing or conflicting entries. The reconciliation method finds one conflicting receipt introduced on purpose. Investigate whether both paths produced downstream actions; an idempotent inference response does not make external side effects idempotent. Preserve disputed decisions for manual handling and avoid a second automated score.

Return with a measured decision

Switch back only after the original region has the current approved digest, fresh features and complete logging. Deliver the recovery timeline, feature lag, time to safe route, duplicate retries, divergent keys, missing logs and operator decision. Link the drill to the incident playbook and revise its hold conditions if the drill exposed a gap. A final green health check without reconciled decisions is not a completed recovery exercise.

Implementation

python
def recovery_state(manifests, decisions):
    primary, standby = manifests
    if primary["digest"] != standby["digest"]:
        return "hold:artifact"
    if standby["feature_lag_seconds"] > 47:
        return "manual-review:lag"
    seen = {}
    for decision in decisions:
        identity = (decision["route"], decision["digest"])
        old = seen.setdefault(decision["key"], identity)
        if old != identity:
            return "hold:divergent-decision"
    return "recover"

manifests = ({"digest": "model-47"},
             {"digest": "model-47", "feature_lag_seconds": 18})
decisions = [{"key": "r-82", "route": "review", "digest": "model-47"}]
assert recovery_state(manifests, decisions) == "recover"
assert recovery_state(manifests, decisions + [{**decisions[0], "route": "clear"}]) == "hold:divergent-decision"

Performance and operating cost

The reconciliation loop is O(n) expected time and O(n) space for n decisions. Duplicate regional capacity and continuous artifact and feature replication are the larger costs. The sample gate is deliberately small; a real drill needs log completeness, caller idempotency and business-action checks before traffic returns.

Common Mistakes

  • Accepting standby traffic before checking its artifact digest.
  • Reusing a stale feature snapshot after a healthy model canary.
  • Counting repeated requests without comparing their issued decisions.
  • Returning to primary while the overlap window still has unresolved log gaps.

Read next

ai-data
mlops
Storage details