Restore metadata and artifacts from different points, reconcile a promotion and prove the live model remains approved.
Project: restore a receipt-model registry without changing live decisions
Create a controlled failure
The receipt-risk service has model r47 live and r46 as rollback. A database backup records r47, but artifact replication for r47 is delayed. A later approved promotion event to r48 exists only in the append-only log, and a draft r49 was registered but never approved. Start the drill by pinning the digest actually served on every cohort. Freeze automated promotion, preserve the deployment pointers and isolate the restore. Backup consistency reveals the missing r47 bytes before any alias is trusted.
Restore and classify
Restore the metadata checkpoint, obtain r47 bytes from the original artifact store or a verified alternate copy, and hash them against the manifest. If the bytes cannot be recovered, keep the loaded r47 replicas and stage a compatible approved fallback rather than restarting. Compare restored alias, r48 promotion event, r49 draft state and actual deployment identity. Reconciliation requires approval and expected previous pointer before r48 can be replayed.
Prove behavior before switching
Load r47 and r48 in the isolated recovery environment, run the frozen receipt smoke set and verify the policy route and output schema. Rebuild the registry alias only after r48 bytes, approval and runtime all pass. Keep r49 dark even if its version number is highest. Check rollback to r47 and record artifact availability in every region that might serve the model. Region failover identity prevents a recovered alias from hiding different deployed bytes.
Measure the recovery result
Publish time to restore metadata, time to obtain verified artifact bytes, time to reconcile promotions and time to re-enable deployment. Report recovery-point loss, service interruption, missing audit records and follow-up repairs. The drill passes only when serving digests match approved pointers and r49 remains unavailable for production. Repeat after registry schema or artifact-store layout changes; a one-time success does not validate a changed recovery path.
Implementation
def disaster_drill_gate(serving, approved, available, smoke_passed):
if not serving:
return "hold:no-serving-evidence"
for cohort, digest in serving.items():
if digest not in approved or digest not in available:
return "hold:" + cohort
if digest not in smoke_passed:
return "hold:smoke-" + cohort
return "restore-complete"
serving = {"north": "receipt-r47", "south": "receipt-r48"}
assert disaster_drill_gate(serving, {"receipt-r47", "receipt-r48"},
{"receipt-r47", "receipt-r48"},
{"receipt-r47", "receipt-r48"}) == "restore-complete"
assert disaster_drill_gate(serving, {"receipt-r47"},
{"receipt-r47", "receipt-r48"},
{"receipt-r47", "receipt-r48"}) == "hold:south"
Performance and operating cost
The final gate is O(c) time and O(1) extra space for c cohorts. The drill itself pays for isolated infrastructure, artifact copies, load tests and operator time. Recovery objectives are meaningful only when measured against the slowest required component, not just database restore duration.
Common Mistakes
- Assuming a metadata restore includes model bytes.
- Promoting the highest restored version without checking approval.
- Restarting loaded production replicas before an artifact copy is verified.
- Reporting database restore time as the total recovery time.
Read next
- Model registry backups: keep metadata and artifact bytes consistent
- Registry restore: reconcile aliases, approvals and serving pointers
- Region failover for inference: match model and feature state
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Model artifacts: verify digest, origin and loading format
