Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: restore a receipt-model registry without changing live decisions

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Restore metadata and artifacts from different points, reconcile a promotion and prove the live model remains approved.

Create a controlled failure

The receipt-risk service has model r47 live and r46 as rollback. A database backup records r47, but artifact replication for r47 is delayed. A later approved promotion event to r48 exists only in the append-only log, and a draft r49 was registered but never approved. Start the drill by pinning the digest actually served on every cohort. Freeze automated promotion, preserve the deployment pointers and isolate the restore. Backup consistency reveals the missing r47 bytes before any alias is trusted.

Restore and classify

Restore the metadata checkpoint, obtain r47 bytes from the original artifact store or a verified alternate copy, and hash them against the manifest. If the bytes cannot be recovered, keep the loaded r47 replicas and stage a compatible approved fallback rather than restarting. Compare restored alias, r48 promotion event, r49 draft state and actual deployment identity. Reconciliation requires approval and expected previous pointer before r48 can be replayed.

Prove behavior before switching

Load r47 and r48 in the isolated recovery environment, run the frozen receipt smoke set and verify the policy route and output schema. Rebuild the registry alias only after r48 bytes, approval and runtime all pass. Keep r49 dark even if its version number is highest. Check rollback to r47 and record artifact availability in every region that might serve the model. Region failover identity prevents a recovered alias from hiding different deployed bytes.

Measure the recovery result

Publish time to restore metadata, time to obtain verified artifact bytes, time to reconcile promotions and time to re-enable deployment. Report recovery-point loss, service interruption, missing audit records and follow-up repairs. The drill passes only when serving digests match approved pointers and r49 remains unavailable for production. Repeat after registry schema or artifact-store layout changes; a one-time success does not validate a changed recovery path.

Implementation

python
def disaster_drill_gate(serving, approved, available, smoke_passed):
    if not serving:
        return "hold:no-serving-evidence"
    for cohort, digest in serving.items():
        if digest not in approved or digest not in available:
            return "hold:" + cohort
        if digest not in smoke_passed:
            return "hold:smoke-" + cohort
    return "restore-complete"

serving = {"north": "receipt-r47", "south": "receipt-r48"}
assert disaster_drill_gate(serving, {"receipt-r47", "receipt-r48"},
                           {"receipt-r47", "receipt-r48"},
                           {"receipt-r47", "receipt-r48"}) == "restore-complete"
assert disaster_drill_gate(serving, {"receipt-r47"},
                           {"receipt-r47", "receipt-r48"},
                           {"receipt-r47", "receipt-r48"}) == "hold:south"

Performance and operating cost

The final gate is O(c) time and O(1) extra space for c cohorts. The drill itself pays for isolated infrastructure, artifact copies, load tests and operator time. Recovery objectives are meaningful only when measured against the slowest required component, not just database restore duration.

Common Mistakes

  • Assuming a metadata restore includes model bytes.
  • Promoting the highest restored version without checking approval.
  • Restarting loaded production replicas before an artifact copy is verified.
  • Reporting database restore time as the total recovery time.

Read next

ai-data
mlops
Storage details