Degraded and crashed are not the same condition

A redundant RAID can continue working after one drive fails. The management interface may report the pool as degraded because it has lost redundancy, not because every file has disappeared.

That can be a routine maintenance event. It can also be the last stable moment before a second problem turns it into data recovery.

The first question is therefore not “where is the rebuild button?” It is:

Is this a healthy, accessible array with one clearly failed member, or an uncertain set that already has missing data, read errors or inconsistent metadata?

The answer determines whether you are maintaining a system or preserving evidence.

When a normal repair is reasonable

The manufacturer’s repair procedure can be appropriate when all of these are true:

  • the volume remains accessible;
  • the RAID level provides enough redundancy for the present failure;
  • one failed member is clearly identified by serial number and slot;
  • the remaining members show no concerning read or hardware errors;
  • the array configuration has not been reset or recreated;
  • a separate backup has been opened and verified;
  • the replacement drive meets the vendor’s requirements.

Synology’s repair guidance describes this expected degraded state: replace the failed drive with a suitable healthy one, then use the storage-pool repair function. It also says that the function does not repair a broken drive or guarantee rescue from a crashed volume.

That last distinction matters. A rebuild is designed to restore redundancy to a known array. It is not designed to discover what the array used to be.

What a rebuild actually does

RAID spreads or mirrors data across several members according to rules: disk order, stripe size, parity rotation, offsets and the RAID level itself.

During a rebuild, the controller reads surviving members, calculates what it believes belongs on the replacement and writes that result. On a large array this can involve reading a substantial part of every remaining disk.

The operation therefore assumes:

  1. the controller has identified the failed member correctly;
  2. the surviving members contain the correct generation of data;
  3. the RAID parameters are correct;
  4. unreadable sectors will not undermine the calculation;
  5. the replacement is the disk intended to receive writes.

If the assumptions are correct, this is ordinary RAID maintenance. If they are wrong, the process can write a convincing but incorrect version of the array.

Stop before rebuilding when any of these applies

The volume is already inaccessible

A degraded but mounted volume and a crashed or unassembled pool are different problems. If shares have vanished, the file system is RAW, the pool refuses to assemble or the interface proposes creating a new volume, do not use creation as a recovery method.

More than one drive is suspect

A RAID 5 set normally tolerates one member failure; RAID 6 normally tolerates two. That arithmetic does not mean a set is healthy whenever the dashboard reports fewer failed drives.

Another member may still be online while producing timeouts, bad sectors or intermittent resets. A rebuild places sustained read demand on the survivors. The drive that fails halfway through may be the one that held the only good copy of a damaged stripe.

Drive order is uncertain

Slot order, serial numbers and controller metadata are evidence. Do not shuffle disks to see which arrangement works. Do not trust a handwritten label without checking the serial number shown by the system.

Someone already started or interrupted a rebuild

Record how far it reached, which disk was used as the target and why it stopped. Do not restart it automatically. The partial write has changed the state of at least one member.

The array was reset, recreated or initialised

Creating a new pool with “the same settings” is not a neutral act. Even when data areas are not immediately overwritten, new metadata can replace information needed to understand the former layout.

The data has no verified backup

RAID is availability, not backup. It can keep a service online after a member failure, but it does not protect against deletion, ransomware, file-system corruption, controller mistakes or a bad rebuild.

Preserve the evidence before changing the set

Take photographs of:

  • the front and rear of the NAS or server;
  • every drive in its original bay;
  • serial numbers and slot labels;
  • the exact warning or storage-pool screen;
  • controller model, firmware and RAID status;
  • recent event-log entries.

Export configuration and logs only if the system provides a read-only or clearly safe route. Do not install packages, update firmware or run a scrub merely to produce more information.

If drives must be removed, label each one with its original bay and keep them together. Do not return an allegedly failed member under warranty until the data has been secured; a drive that drops from the array can still contain valuable stripes or older metadata.

Should you copy files while the array is still online?

If the volume is stable, quiet and readable, a verified backup of irreplaceable files is the priority. Start with the data that cannot be recreated, not a full copy of software installers or replaceable media.

But a live copy is still a heavy read operation. Stop if:

  • another drive starts clicking or dropping offline;
  • read errors rise;
  • the NAS freezes or repeatedly restarts;
  • the controller marks another member critical;
  • the only-copy data is valuable enough that uncontrolled failure is unacceptable.

At that point, continuing a file-level copy can be less safe than controlled imaging of the individual members.

How professional RAID recovery differs from a rebuild

The aim is to avoid modifying the only sources.

Each member is identified and assessed separately. Unstable disks are imaged using strategies designed for damaged media rather than treated as ordinary healthy storage. DeepSpar describes multi-pass imaging, head-specific handling and controlled responses to timeouts; these are examples of acquisition controls, not guarantees for a particular array.

The RAID is then reconstructed virtually from the member images. The engineer tests possible order, parity, stripe and offset against file-system evidence. Recovered files are written to separate storage.

Professional suites such as the RAID editions listed in ACE Lab’s product catalogue support multiple standard and custom array types. The important principle is not the product name: source members are evidence, and reconstruction should not require writing assumptions back onto them.

A decision table for the first hour

What you see Sensible next step
One failed drive, volume accessible, verified backup Follow the exact vendor repair procedure with a compatible replacement.
One failed drive, no backup, remaining drives healthy Back up the most important data before rebuilding.
Volume inaccessible or pool shown as crashed Stop services and preserve every member; do not initialise or create.
Two or more suspect drives Do not rebuild on the original set; arrange member-level diagnosis.
Drive order changed Stop and reconstruct the original order from serials, bays, logs and metadata.
Rebuild already failed Record the target drive and progress; keep all original and replacement members.
Ransomware or mass deletion Isolate the system; rebuilding redundancy will not restore earlier file content.

The table cannot diagnose an array remotely, but it prevents the most common category error: using a maintenance operation to solve a recovery problem.

After the incident

Once the data is safe, improve the system around the array:

  • keep an offline or immutable backup;
  • test a restore, not only the backup job;
  • configure alerts to reach a person who will act;
  • retain a map of chassis, slot and serial number;
  • record the RAID level and controller configuration;
  • replace ageing members under a planned policy;
  • rehearse what happens when a pool becomes degraded.

A hot spare can reduce the time spent without redundancy, but an automatic rebuild still reads the remaining members and does not replace a backup.

If the array is inaccessible or another disk is unstable, leave every member as it is and request a RAID data recovery diagnosis. The best chance usually begins before somebody clicks the button that promises to make the warning disappear.

Questions people ask next

Does a degraded RAID mean my data is already lost?
Not necessarily. A redundant array can remain accessible after one member fails, which is the purpose of its redundancy. The risk is that protection has been reduced and the remaining drives must carry the workload, so backup verification and accurate fault identification become urgent.
Should I replace the failed NAS drive and click Repair?
Only when the array is genuinely degraded rather than crashed, the failed member is unambiguous, the other members are healthy and a verified backup exists. Vendor repair is a maintenance procedure for a known state, not a safe experiment for an inaccessible or partly reconstructed volume.
Why can a RAID rebuild damage recovery chances?
A rebuild writes calculated data to a member using the controller’s current view of disk order, parity, stripe size and failure state. If that view is wrong, or another member returns bad data, the operation can replace useful evidence with a coherent but incorrect reconstruction.
Can I move all drives into another NAS enclosure?
Do not do so casually. A compatible replacement chassis may import an existing configuration, but another system may offer to initialise or migrate it and can interpret slot order or metadata differently. Record the original arrangement and seek model-specific guidance before moving members.