A registry retention rule can delete an image that is no longer tagged even while a deployment manifest still references its digest. Running Pods may keep serving because the image is cached on their nodes, yet a replacement Pod or rollback cannot pull the deleted object. The retention decision must be based on deployed, rollback, and recovery-region references rather than tag age alone.
Rollback image retention: keep every approved fallback pullable
Operational decision
A settlement API keeps two approved versions for rollback. Before a registry lifecycle change, export the exact digests referenced by live workloads, pending rollouts, and recovery manifests. Compare them with candidate expirations and retain both current and rollback digests in every required region. The text fragment describes the release inventory that a cleanup job must consult; its values are synthetic. Test by replacing one Pod after the old digest would otherwise have expired, and verify it pulls from the registry rather than relying on a node cache. Then execute a rollback in a disposable environment and check a real settlement request. Mark a digest eligible for deletion only after its rollback window closes, no workload or recovery plan references it, and a newer approved fallback is verified. If the registry supports tag-based retention, keep protected tags as a convenience but still compare actual digest references.
Settlement retention gate
Current digest: sha256:47a4...
Rollback digest: sha256:82bd...
Recovery region: both digests present and pull-tested
Delete candidate: absent from live, rollout, and rollback manifests
Proof: replacement Pod pulls and synthetic settlement succeedsCost and verification
Longer retention uses storage and may increase inventory-review time. Overaggressive cleanup reduces spend briefly but can turn a routine restart into an outage. A successful running Pod is weak evidence because its node may have cached layers. Observe registry pull errors, missing digest checks, and rollback drill outcomes. Cleanup should be reversible where possible and should never infer safety from the absence of a human-friendly tag alone.
Common Mistakes
- Do not delete an untagged image without checking digest references.
- Do not call a cached running Pod proof that the registry can still serve the image.
- Do not expire the prior release before the new one passes recovery-region checks.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Immutable artifacts and release provenance
- Registry replicas: prove the recovery region has the exact release image
- Progressive delivery: canary checks and rollback
- Release evidence: tie one deployed digest to one approval decision
