RAID is a way of arranging multiple drives to achieve a particular balance of capacity, performance, and tolerance for drive failure. It is not a backup strategy and it does not make every storage workload faster. The decision criteria in the current 220-1201 Core 1 make more sense when RAID is treated as one layer inside a complete storage design.
The practical foundations in RAID architecture and modern storage technologies point to the same idea: the right layout depends on the workload and on what the organization is trying to survive. A two-drive mirror, a striped performance volume, a parity array, and a standalone NVMe SSD solve different problems.
Start with hard requirements: usable capacity, write and read pattern, tolerated downtime, number of drive failures to survive, rebuild time, controller/platform support, and recovery requirements. Then choose a layout. Starting with a RAID number and inventing a justification afterward reverses the engineering process.
RAID 0 trades resilience for parallelism and capacity
Striping data across multiple drives can improve throughput and combines their capacity, but the array depends on every member. Lose one drive and the striped dataset is lost.
That can be acceptable for temporary scratch data, caches, or workloads whose authoritative copy exists elsewhere and can be rebuilt. It is usually a poor place for the only copy of user or business data.
RAID 0 demonstrates why ‘more disks’ is not automatically ‘more reliable.’ The mechanism determines the failure domain.
A striped scratch array should still be monitored. Losing a member may be acceptable from a data-recovery perspective, but a sudden failure during a long render can waste hours of work. ‘Data is reproducible’ does not mean downtime or rerun cost is irrelevant.
RAID 0 should also be evaluated against recovery time. If recreating the scratch dataset takes eight hours, losing the array may be acceptable for data integrity but unacceptable for project schedule. Availability decisions should include the cost of regeneration as well as whether data is theoretically replaceable.
RAID 1 spends capacity on a second copy
A two-drive mirror stores the same data on both members. One drive can fail and the volume can continue from the other, assuming the platform and remaining hardware behave normally.
Usable capacity is roughly one drive’s capacity in the common two-drive case. Reads may benefit from multiple members depending on implementation, while writes must be committed to both copies.
A mirror is simple and recoverable, but it still replicates file deletion, ransomware encryption, corruption, and operator mistakes. Redundancy is not versioned recovery.
Mirrors also need attention after a drive failure. The system may continue normally enough that nobody notices the degraded state. Alerts, spare-part availability, and replacement procedures determine how long the array remains one failure away from total loss.
Mirrored arrays can hide latent errors until rebuild. If the surviving disk contains unreadable sectors, the replacement process may expose a second problem. Regular health checks and backups matter because a mirror preserves availability after one failure but cannot guarantee the remaining copy is perfect.
RAID 5 and parity designs change write behavior
RAID 5 distributes data and parity across at least three drives so the array can survive one drive failure. The parity calculation and additional I/O required for small writes can make write performance different from a simple stripe or mirror.
Large arrays with slow drives can also spend significant time rebuilding after a failure. During rebuild the remaining drives are under load and the array has reduced fault tolerance.
Choose parity because capacity efficiency and fault tolerance fit the workload, not because it seems like the default middle option.
Parity arrays are especially sensitive to write pattern and controller behavior. Small random writes can require reading old data/parity and writing updated data/parity, while large sequential operations behave differently. Benchmark the expected workload instead of extrapolating from a simple sequential speed test.
Parity layouts also complicate controller-cache requirements. Write-back caching can improve small-write performance, but unsafe cache without battery or flash protection can increase corruption risk during power loss. Performance features should be evaluated with durability behavior, not benchmarked in isolation.
RAID 10 uses more disks to simplify failure recovery
RAID 10 combines mirroring and striping, commonly requiring at least four drives. It sacrifices more raw capacity than single-parity layouts but can provide strong read/write performance and simpler recovery characteristics.
The exact failures it can survive depend on which members fail. Losing both sides of the same mirror pair can be fatal even when the total number of failed drives is small.
The correct comparison therefore includes layout topology and rebuild behavior, not just the phrase ‘can survive two disks’ repeated without conditions.
RAID 10 can rebuild from a surviving mirror copy rather than recalculating parity across many drives, which can be operationally attractive. But the higher raw-capacity cost and minimum drive count may be unjustified for small systems whose recovery objective can be met with simpler storage plus backup.
Drive type changes the economics of the array
SATA HDDs, SATA SSDs, and NVMe SSDs have very different latency, throughput, endurance, and interface behavior. The evolution from SATA to NVMe means an array built to hide hard-disk latency may not have the same value on low-latency solid-state storage.
RAID controllers and software stacks can also become bottlenecks before the drives do. Verify interface lanes, controller bandwidth, cache behavior, and platform support.
Do not buy premium NVMe drives for a workload constrained by a 1 Gb/s network share and expect users to see the media’s benchmark throughput.
Drive endurance and failure correlation matter too. Buying identical drives from the same batch can create similar age and wear profiles. Solid-state arrays also have write-endurance limits. Redundancy handles individual device failure; lifecycle planning handles fleets of devices aging together.
Rebuild time is an availability requirement
A failed disk begins a degraded period. The organization must detect the failure, replace the member, and rebuild or resilver the array before it returns to normal redundancy.
Rebuild duration depends on array size, drive performance, workload activity, controller behavior, and how much data must be reconstructed. Larger drives can increase the time spent exposed to a second failure.
Monitor array health and test alerting. Redundancy has little value if the first failed member goes unnoticed for months.
Rebuild operations can reduce application performance dramatically. A system may technically remain available while user experience violates the service objective. Decide whether rebuild should run at maximum speed, be rate-limited, or trigger workload movement depending on the business impact.
Hot spares can shorten time-to-rebuild by allowing recovery to begin automatically after a failure, but they consume capacity and do not replace monitoring. An automatic rebuild that nobody notices can still leave the array degraded again when another member fails later.
RAID does not replace backup
RAID can keep a volume available after some hardware failures. Backup creates an independent recovery copy for deletion, corruption, ransomware, catastrophic controller failure, theft, disaster, or mistakes propagated across all members.
The general fault-tolerance principle is that surviving one fault requires independent controls around different failure modes. A mirror protects against one disk failing; an offline or remote backup protects against a broader set of data-loss events.
Define recovery point and recovery time separately from array availability. If the business needs last week’s clean version, no RAID level can provide it by itself.
Backup testing should include bare restore or file restore from outside the array. If the only backup catalog, encryption key, or recovery software lives on the failed storage system, the independent copy is less useful than the organization assumes.
A decision example makes the trade-offs concrete
A small design workstation with two SSD bays and daily versioned backup may choose RAID 1 to reduce downtime from one drive failure. A scratch-render workstation might prefer RAID 0 because the source assets live elsewhere. A four-drive application host may use RAID 10 when write performance and rapid degraded operation matter more than raw capacity.
None of those is a universal winner. Each follows a different consequence of failure.
Also consider the alternative of one reliable SSD plus excellent backup and rapid replacement. Additional array complexity is justified only when local availability or throughput actually needs it.
Controller failure can be another shared dependency. Hardware RAID metadata, cache, battery/flash protection, and proprietary layouts may affect how easily disks can be moved to another controller. Record the controller model and recovery requirements instead of assuming the disks alone contain everything needed.
Choose storage as a recovery system
Document usable capacity, expected workload, drive type, tolerated failures, alerting, spare/replacement process, rebuild expectations, controller dependency, and backup/restore plan.
Test that operators can identify a failed member and restore from backup. A healthy array dashboard is not proof that the recovery copy works.
The CompTIA A+ certification perspective is practical: understand RAID levels, but choose and troubleshoot them as part of the whole storage path. The best layout is the one whose performance, failure behavior, and recovery process match the system’s actual requirements.
Storage decisions should be revisited when workload changes. A file server that began as mostly reads can become a write-heavy virtual-machine datastore; a scratch array can become the unofficial location for irreplaceable user data. The original RAID choice is only valid while the assumptions behind it remain true.
A storage decision record should note why RAID was selected over simpler alternatives. If local availability is the reason, measure downtime avoided. If throughput is the reason, benchmark the application. If neither benefit appears in practice, the array may be adding complexity without delivering value.