A Kubernetes VolumeSnapshot asks a compatible storage driver to capture a PersistentVolumeClaim's data. It does not by itself guarantee a transactionally consistent database backup, copy the application's configuration, or prove that another cluster can restore the data. The snapshot controller and CSI driver must support the operation for the chosen storage class.
Volume snapshots: test application-consistent restore
Operational decision
A ledger service stores its database on a PVC. Quiesce writes or use the database's supported backup protocol before capturing a volume snapshot; record the database position, schema version, image digest, and any encryption-key dependency. The manifest requests a snapshot in the same namespace as the source claim. Restore it to a new PVC in a disposable namespace, start the compatible database version, and run a synthetic balance check. Measure elapsed restore time and recovered transaction position against the recovery objectives. A snapshot kept in the same failure domain can vanish with the storage system; copy or export recovery data according to the failure model. Keep independent backups for bad writes that were already reflected in recent snapshots.
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata: {name: ledger-checkpoint-47, namespace: ledger}
spec:
volumeSnapshotClassName: ledger-csi-snapshots
source:
persistentVolumeClaimName: ledger-dataCost and verification
Snapshots consume storage and can incur transfer cost when copied across regions. Fast capture often shifts work to restore time, when blocks must be rehydrated before the database meets its latency target. A successful snapshot status proves a storage operation, not application integrity. Test restores with realistic data size and key access. Track retention and deletion separately from the source PVC so routine cleanup cannot erase the only usable recovery point.
Common Mistakes
- Do not call a crash-consistent disk image a verified database backup.
- Do not keep every recovery point in the same failure domain.
- Do not skip the restore drill after a successful snapshot event.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Persistent storage: PVC lifecycle and data ownership
- Backups and disaster recovery: prove the restore path
- Database change safety: expand, migrate, contract
- Multi-region failover: define write ownership before moving traffic
Practice and check
Advanced follow-up
- PV reclaim policy: trace the real asset before deleting a claim
- Raw block PVCs: make the application's formatting responsibility explicit
