Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Linux md RAID recovery: identify the failed member before rebuilding

Last updated: 5 Oct 20267 min read
tutorial
AdvancedBy AITrove Editorial

Software RAID status identifies array level, active members, failed members, and rebuild progress. The member path alone is not enough to locate physical hardware or the provider attachment: map it through stable device IDs, serials, and array metadata before touching a disk. Adding a replacement to a degraded array can start reconstruction, which competes with application I/O and exposes unreadable blocks on surviving members. RAID is not a backup; deletion, corruption, and many operator mistakes propagate across members. Recovery therefore protects a recent independent backup, limits application load if needed, and checks both the array and data after rebuild.

Operational decision

A four-member RAID10 array backing an image cache reports one failed member. The responder records md status, event count, member serials, and the last successful restore test, then confirms which physical or virtual attachment failed through an independent inventory. A disposable array exercise removes that member, observes service latency, and adds a known blank replacement only after matching its serial to the intended slot. During rebuild, the team tracks progress, read errors, and API tail latency. A second fault simulates a mistaken replacement selection; the runbook must stop before writing to a healthy member. Return to service requires a nondegraded array state and a sample of image hashes read through the application.

bash
cat /proc/mdstat
mdadm --detail /dev/md0
lsblk -o NAME,SERIAL,SIZE,TYPE

Cost and verification

Reconstruction consumes disk bandwidth and can lengthen client response times; an aggressive speed target may deepen latency, while a very slow target prolongs reduced redundancy. Monitor both rebuild state and user-path errors. A completed rebuild does not certify the newest application backup or catch corruption that predates the member loss. Keep evidence of the member identity, replacement mapping, and post-rebuild checks. If the array level has no redundancy, do not describe a failed member as a routine degraded rebuild: recovery may require restoring from backup.

Common Mistakes

  • Do not remove a member based only on a transient device letter.
  • Do not treat RAID redundancy as an application backup.
  • Do not declare recovery complete at 100% rebuild without data and service checks.

Connected lessons

Practice and check

devops
linux
storage-operations
Storage details