Cross-region registry replication is asynchronous in many setups. A rule configured in the source registry does not prove a newly pushed release has arrived in the destination. Repository permissions, lifecycle settings, encryption behavior, and tag rules may differ across registries. A recovery cluster also needs working pull credentials and network access to the destination, not just stored image bytes.
Registry replicas: prove the recovery region has the exact release image
Operational decision
A document parser is deployed in one region and kept ready for regional recovery. For each approved release, record the source digest and every platform-specific child digest if the image is multi-architecture. Query the recovery registry for the expected object and verify its manifest identity before declaring the release recoverable. The sample is read-only and inspects one digest in the destination account. In a disposable recovery cluster, pull that digest with the actual workload identity, start the container on each target architecture, and run a synthetic parse. Check destination repository policy, encryption key access, and retention independently; do not assume source settings replicated. Delay one image copy deliberately and ensure the failover gate blocks promotion to the missing digest. Retain the previous known-good release in both regions until the new one passes the target-cluster test.
aws ecr describe-images --repository-name document-parser --image-ids imageDigest=sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa --region ap-south-2 --query 'imageDetails[].{digest:imageDigest,tags:imageTags}'Cost and verification
Replication adds registry storage and transfer cost. Holding two known-good releases adds retention cost but makes rollback possible when the newest image is missing or broken. A digest lookup proves registry presence, not that a Pod can authenticate, pull, and run the image. Track replication age, missing digests, pull failures, and synthetic transaction results per region. A fallback to a mutable tag can silently deploy different bytes.
Common Mistakes
- Do not treat a configured replication rule as proof of an individual digest.
- Do not omit destination pull credentials and network access from a failover test.
- Do not assume repository policies and retention settings follow the image copy.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Multi-region failover: define write ownership before moving traffic
- Object replication: verify the exact recovery object arrived
- Multi-architecture images: verify every platform behind one tag
- Registry tag mutation: reject release decisions based on a moving pointer
