Build a disposable Linux storage lab with a synthetic ledger service, a separate data mount, a four-member RAID10 array on virtual disks, and an independent backup. Keep a ledger of device identities, filesystem types, mount sources, application sequence IDs, and measured request latency. Save a pre-fault data copy. No repair, resize, or member replacement in this exercise should touch an irreplaceable volume.
Linux storage fault drill: space, mount, repair, RAID, and latency
Trace pressure and layered growth
Fill the ledger mount with controlled test data, then unlink a large open log. Compare df, du, inode counts, and open-file evidence; free space must return only after the owner releases the descriptor through its supported path. Expand a clone through its actual provider, partition, encrypted mapping, logical volume, and filesystem layers as applicable. Record each layer's before-and-after size and prove a synthetic write at the mounted path.
Fail the mount and filesystem
Boot a VM without its expected data volume. The management plane should remain reachable while the ledger writer stays stopped and no fallback files appear under the bare mountpoint. Inject a read-only mount or simulated device error on a clone. Preserve the first kernel and application errors, stop writes, and practice type-specific inspection or repair only after the filesystem is detached. Compare recovered ledger sequence IDs with the independent backup and application records before admitting a writer.
Acceptance ledger
Space: source and owner identified; write succeeds after controlled release
Growth: provider-to-filesystem layer sizes agree at the mounted path
Missing mount: writer stopped; bare mountpoint remains empty
Repair: detached clone checked; application sequence IDs reconciled
RAID: replacement serial verified; array nondegraded; hashes match
Latency: device and application intervals aligned; throughput recordedRecover redundancy and diagnose delay
Remove one member from the disposable redundant array, map the failed slot by serial and metadata, then rebuild onto a known replacement. Track rebuild progress, unreadable blocks, and user-path latency. Refuse a deliberately mislabeled replacement before it is added. Finally, apply a controlled write load, compare per-interval device waits with application queue time, and change only one concurrency limit. The final evidence includes the array state, file hashes, request outcomes, and a repeatable rollback to the original load setting.
Common Mistakes
- Do not repair a mounted production filesystem during the drill.
- Do not identify a replacement disk by a changing device letter.
- Do not declare a storage change complete without a real application write and read.
Connected lessons
- Linux disk pressure: explain missing space before deleting application data
- Linux filesystem expansion: follow the block device to the mounted filesystem
- Linux read-only filesystem incidents: preserve evidence before repair
- Linux mount boot contracts: prevent a service from writing into an empty mountpoint
- Linux md RAID recovery: identify the failed member before rebuilding
- Linux I/O latency: distinguish device wait from an application queue
- DevOps projects
