It is indeed very interesting. They are one of the few companies in the world that are able to perform this kind of analysis. However I note they area using cheap commodity ATA drives I wonder if the failure rate for Fibre Channel ones would follow the same patterns?
I just had a drive fail in a fiber RAID array (hey, “I” still means inexpensive, right?) last week. Making preparations to test a major application upgrade, so customer says to recreate a test db on the standby machine, which is in a remote location with no technical admin staff on site. OK, no problem, done it before, have scripts ready. Bring standby up read-only, do an export, takes a few hours, no problem.
As the export progresses, customer gets call that an alarm is going off, users are very annoyed, and incapable of saying exactly which device (of many in racks) is complaining. Customer asks me to check standby machine for warnings (hp-ux), no warnings. Eventually through a combination of the raid admin software and users saying which lights are flashing, we figure out the array lost a disk drive, alarm means no more hot spare disks. So customer tells them to take out the disk with the flashing light. Which they do. Then they yank out the disks on either end of the array.
they area using cheap commodity ATA drives I wonder if the failure rate for Fibre Channel ones would follow the same patterns
Jason, see StorageMojo post on that – http://storagemojo.com/?p=383
It is indeed very interesting. They are one of the few companies in the world that are able to perform this kind of analysis. However I note they area using cheap commodity ATA drives I wonder if the failure rate for Fibre Channel ones would follow the same patterns?
jason.
I just had a drive fail in a fiber RAID array (hey, “I” still means inexpensive, right?) last week. Making preparations to test a major application upgrade, so customer says to recreate a test db on the standby machine, which is in a remote location with no technical admin staff on site. OK, no problem, done it before, have scripts ready. Bring standby up read-only, do an export, takes a few hours, no problem.
As the export progresses, customer gets call that an alarm is going off, users are very annoyed, and incapable of saying exactly which device (of many in racks) is complaining. Customer asks me to check standby machine for warnings (hp-ux), no warnings. Eventually through a combination of the raid admin software and users saying which lights are flashing, we figure out the array lost a disk drive, alarm means no more hot spare disks. So customer tells them to take out the disk with the flashing light. Which they do. Then they yank out the disks on either end of the array.
they area using cheap commodity ATA drives I wonder if the failure rate for Fibre Channel ones would follow the same patterns
Jason, see StorageMojo post on that – http://storagemojo.com/?p=383
I was going to mention that too.
It would appear I’d better be quick round here!