Preventing RAID problems

By Mitch Tulloch, ITworld.com |  Operating Systems, array, drive Add a new comment

It's always better to prevent problems than to try and resolve them when they occur. And one or the primary failure points of a server (or a high-powered workstation) is the storage subsystem. This is because hard drives are mechanical doohickeys that are prone to sudden failure, with either accompanying operating system failure, data loss, or both.

Using RAID storage is supposed to help anticipate such catastrophes by providing a level of fault tolerance for your storage subsystem. But even RAID can have its problems, especially when you use some newer high-capacity hard drives in the near-terabyte range that have less than stellar reputations for reliability. A colleague found this out recently when one of the drives in his RAID 6 array failed after only a year of use. Since RAID 6 provides an extra level of redundancy over RAID 5 and can survive the loss of two drives without data loss, my colleague felt it was safe to contact the manufacturer and request a replacement drive for the one that failed.

Fortunately, by the time the replacement had arrived, no other drive had failed. My colleague then removed what he thought was the failed drive from the RAID array, only to discover he had removed a working drive instead. Now if he had been running RAID 5 at that point, he'd be toast. Fortunately, he was able to resolve the situation by (a) plugging the working drive he had removed back into the array (b) waiting for the array to rebuild the parity info (c) removing the actual failed drive and replacing it with the replacement from the manufacturer and (d) waiting for the rebuild of the second level parity info. Everything worked fine, but he got to bed quite late that night.

What's the solution to avoiding near-disasters like this? Simple: label the drives in your RAID array so you can match them to the ports on your controller card and to drive numbers in your controller's management software. And while you're at it, you may as well label all your drive cables as well, and your drive bays also. A little preventative maintenance like this can be a big time-saver by preventing you from doing something stupid when the heat is on.

ITworld LIVE

Operating SystemsWhite Papers & Webcasts

White Paper

A Comparison of PowerVM and VMware vSphere (4.1 & 5.0) Virtualization Performance

This technical white paper presents benchmark results showing greater VM consolidation ratios than demonstrated in previous benchmarks and demonstrating the extent of the performance lead that PowerVM virtualization technologies deliver over x86-based add-on virtualization products.

White Paper

Consolidating Lotus Domino x86 Workloads on IBM Power Systems

Read the white paper to learn how moving up to Lotus Domino 8.5 and consolidating with IBM Power Servers can help you boost performance results and ROI.

White Paper

Task, workflow & issue management for teams. Try free!

Need a flexible system for managing team tasks, issue tracking, and automating and managing workflow processes? Comindware® Tracker helps you do it all.

Webcast On Demand

Best Practices in Monitoring VMware

The benefits of virtualization are unassailable: increased agility, scale, and cost savings to name a few. However, so too are the monitoring challenges posed by these environments-including complexities, lack of visibility and control, and inefficiency.

Sponsor: Nimsoft

White Paper

How Nimsoft Service Desk Speeds Deployment and Time to Value

For years, many support teams have been hamstrung by their traditional service desk platforms, which require complex, time-consuming coding for virtually every aspect of customization. This complexity makes it costly and difficult for support organizations to adapt-and places an increasingly substantial burden on the agility and efficiency of the business as a whole.

See more White Papers | Webcasts

Ask a question

Ask a Question