Software RAID status identifies array level, active members, failed members, and rebuild progress. The member path alone is not enough to locate physical hardware or the provider attachment: map it through stable device IDs, serials, and array metadata before touching a disk. Adding a replacement to a degraded array can start reconstruction, which competes with application I/O and exposes unreadable blocks on surviving members. RAID is not a backup; deletion, corruption, and many operator mistakes propagate across members. Recovery therefore protects a recent independent backup, limits application load if needed, and checks both the array and data after rebuild.
Linux md RAID recovery: identify the failed member before rebuilding
Operational decision
A four-member RAID10 array backing an image cache reports one failed member. The responder records md status, event count, member serials, and the last successful restore test, then confirms which physical or virtual attachment failed through an independent inventory. A disposable array exercise removes that member, observes service latency, and adds a known blank replacement only after matching its serial to the intended slot. During rebuild, the team tracks progress, read errors, and API tail latency. A second fault simulates a mistaken replacement selection; the runbook must stop before writing to a healthy member. Return to service requires a nondegraded array state and a sample of image hashes read through the application.
cat /proc/mdstat
mdadm --detail /dev/md0
lsblk -o NAME,SERIAL,SIZE,TYPECost and verification
Reconstruction consumes disk bandwidth and can lengthen client response times; an aggressive speed target may deepen latency, while a very slow target prolongs reduced redundancy. Monitor both rebuild state and user-path errors. A completed rebuild does not certify the newest application backup or catch corruption that predates the member loss. Keep evidence of the member identity, replacement mapping, and post-rebuild checks. If the array level has no redundancy, do not describe a failed member as a routine degraded rebuild: recovery may require restoring from backup.
Common Mistakes
- Do not remove a member based only on a transient device letter.
- Do not treat RAID redundancy as an application backup.
- Do not declare recovery complete at 100% rebuild without data and service checks.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Backups and disaster recovery: prove the restore path
- Volume snapshots: test application-consistent restore
- Capacity and load tests: identify the next bottleneck
- Linux disk pressure: explain missing space before deleting application data
- Linux filesystem expansion: follow the block device to the mounted filesystem
