A local PersistentVolume represents storage tied to a specific node and carries node affinity so the scheduler can place a consuming Pod there. It can provide low-latency I/O, but another node cannot attach the same bytes when the original host fails. Pod rescheduling and storage recovery are separate operations. A healthy replica count before failure does not establish that data on a lost local volume can be served elsewhere.
Local PVs: plan for node loss as a data-location failure
Operational decision
A receipt index stores rebuildable search segments on local SSD. Map each PV to node, zone, ordinal, and source data checkpoint. Use a StorageClass with delayed binding so the first consumer and available local PV are selected together; inspect the bound PV's node affinity before launch. Kill the test node and observe that the affected Pod cannot simply use that same local volume on a replacement node. Rebuild its segment from a verified upstream checkpoint onto a new claim, then compare document count and query results before rejoining traffic. If the local data is not rebuildable, the design needs replication or a backup and restore procedure with a measured recovery objective before deployment. During node maintenance, drain only after the replacement data path is ready; moving the Pod object alone does not move the disk. Track abandoned PVs on dead hosts for secure disposal and accounting.
Receipt local-volume recovery record
PV identity: claim, node, zone, ordinal
Data source: upstream checkpoint and generation
Node loss: old disk unavailable to replacement node
Recovery: create new claim, rebuild, compare count and queries
Traffic gate: rejoin only after data verification
Cleanup: retire old PV and physical media under owner approvalCost and verification
Local SSD can improve latency and avoid network-storage charges, but it moves availability cost into replication, rebuild traffic, spare capacity, and operations. A large index may take hours to reconstruct, so calculate recovery time from actual ingest rate and throttling limits. Measure rebuild duration, source load, missing documents, and PVs stranded on failed nodes. Do not present a local PV as interchangeable with a network block volume during incident response.
Common Mistakes
- Do not expect a local PV to attach to a replacement node.
- Do not route queries to a rebuilt shard before its generation is verified.
- Do not choose local storage for irreplaceable data without a tested recovery route.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Persistent storage: PVC lifecycle and data ownership
- Delayed volume binding: choose storage topology with the first Pod
- Node autoscaling: make pending Pods schedulable before traffic rises
- Backups and disaster recovery: prove the restore path
