Whats interesting for me is that they have almost exactly 10 time the amount drives we have in production (excluding system drives)
On average we get around 4 failures a week(some weeks more, others none) They tend to fail in bunches. For us there is a correlation between the way they are used and failure rate.
One bunch of file servers was being used for particle sims to model black holes. We pulled petabytes through the array. within 2 months of the show finishing we had replaced about 30% of the drives. (raid 6 with 14 disk luns, 4 hot spares).
Other have been happy and not killed as many drives. yet they've had the same amount of data pulled through them. They are also identical. They have the same drives, raid controller and bought very close together.
One thing to note is that if your work load changes, so will your failure pattern.
Is it possible that the drives that died were in similar physical locations? For example, if you have a dodgy power supply, perhaps the last bank of drives gets lower quality power, or depending on the temp of the case and airflow, perhaps the top bank of drives run significantly warmer?
Thats certainly possible, although we've done our best to minmise it.
We have coldlogik water cooled racks which almost eliminate hot spots. The other thing to note is that 4000 drives fit into less than 7 racks. (yup, I was surprised too) so they are all in the same place.
We also have two transformers, just for us, with lots of sexy power smoothing (but no UPS, and yes that's a bad thing.)
It could be the raid enclosure its self, that might be part of the problem. However they should be identical, with the same firmware.
I'd like to know what sort of failures they are, since that could point to the cause; do they spin up and click repeatedly, refuse to spin up at all, or spin up fine but remain unresponsive?
Thats something the raid handles. Plus its loud in the server room, we have to wear ear defenders. I'd assume that in most cases its on smart errors. However, as I let Dell do all the heavy lifting its pure speculation.
From talking to ex-seagate people most errors are because of misalignment, X-IO (http://xiostorage.com/) make their money by bundling Re-alignment tooling into their disk packs. If I was in the market for pure block storage again, these are the guys I'd buy. (they call it remanufacture)
On average we get around 4 failures a week(some weeks more, others none) They tend to fail in bunches. For us there is a correlation between the way they are used and failure rate.
One bunch of file servers was being used for particle sims to model black holes. We pulled petabytes through the array. within 2 months of the show finishing we had replaced about 30% of the drives. (raid 6 with 14 disk luns, 4 hot spares).
Other have been happy and not killed as many drives. yet they've had the same amount of data pulled through them. They are also identical. They have the same drives, raid controller and bought very close together.
One thing to note is that if your work load changes, so will your failure pattern.