Any DELL or RAID experts around?
Discussion
Looks like the aircon in our server room failed, causing all the servers to overheat and power off. I've got back 8 of the 9 DELLs running fine, but one of them is reporting "unknown" status for the two RAID5 containers.
I'm currently running a verify media check from the RAID controller utils in the hopes that this will fix it, but if not, I'm well and truly FUBAR'd and will have to restore from a backup onto the standby server.
In the event that the verify doesn't help, does anyone have any other suggestions? I know it's a long shot at lunchtime on the hottest Sunday for a long while, but hey, I'm desperate!
Thanks for any help!
I'm currently running a verify media check from the RAID controller utils in the hopes that this will fix it, but if not, I'm well and truly FUBAR'd and will have to restore from a backup onto the standby server.
In the event that the verify doesn't help, does anyone have any other suggestions? I know it's a long shot at lunchtime on the hottest Sunday for a long while, but hey, I'm desperate!
Thanks for any help!
That's why I never EVER run anything except Raid 10 on anything approaching critical.
Raid5 is just a cheap way of getting 'some' redundancy, but its slow and not really that useful except in environments where you don't need to backup data, but then why raid!
Only time I would advocate Raid1 is if you only have 3 drive bays, and cant afford to mirror from a space perspective.
Still, not fun that it happened to you ;(
J
Raid5 is just a cheap way of getting 'some' redundancy, but its slow and not really that useful except in environments where you don't need to backup data, but then why raid!
Only time I would advocate Raid1 is if you only have 3 drive bays, and cant afford to mirror from a space perspective.
Still, not fun that it happened to you ;(
J
Only 3 Drive bays in a poweredge 16xx so I thought I was being clever specifiying RAID5 when I ordered the servers.
Who knew two disks would go at once! If there was a hot spare it would have survived but there isn't so it didn't.
Oh well, c'est la vie, the restore is running and I'm back home eating a sandwhich, will go finish up later.
As for reducing the shutdown temp - I thought I had all that covered but I guess not! Once things are back up and running that's job 1.
Who knew two disks would go at once! If there was a hot spare it would have survived but there isn't so it didn't.
Oh well, c'est la vie, the restore is running and I'm back home eating a sandwhich, will go finish up later.
As for reducing the shutdown temp - I thought I had all that covered but I guess not! Once things are back up and running that's job 1.
If you've only 3 drives, and dont NEED 2x Disks worth of space (ie can live on 1x disk space) then just use Raid 1 with a Hot Spare.
Theoretically can lose 2 disks over time and still deliver 100% performance.
Raid5 is only useful if you NEED 2x Disks worth of space.
If you dont, its just unneccesary performance loss / risk over Raid 1/10
J
Theoretically can lose 2 disks over time and still deliver 100% performance.
Raid5 is only useful if you NEED 2x Disks worth of space.
If you dont, its just unneccesary performance loss / risk over Raid 1/10
J
It was specced originally with 3x 17GB disks, with the idea that RAID5 gave some degree of redundancy - a disk can easily be replaced.
oops, maybe not. In hindsight I probably should have specced it with 3x36GB disks and had RAID1 with a hot spare, but who would predict two disks going at once!
That makes three failed disks in as many months here - time to send a job lot off to seagate for replacement!
Given the sizes and age, they are SCSI disks.. In my experience, the only scsi drives that die are hot scsi disks (or VERY Old > 6 yrs) I think you REALLY need to investigate better cooling of the servers. This is why people spend tens of thousands a year to have their servers in our datacentre suites.
I would take 10 + Hotspare over 5 any day of the week, Disks are cheap compared to a failed RaidSet.
I personally would never use Raid5, its just wrong.
If you're replacing these 2 now, I would go the whole hog and replace all 3 with bigger ones and go 10... you wont regret it
J
I would take 10 + Hotspare over 5 any day of the week, Disks are cheap compared to a failed RaidSet.
I personally would never use Raid5, its just wrong.
If you're replacing these 2 now, I would go the whole hog and replace all 3 with bigger ones and go 10... you wont regret it
J
agent006 said:
_dobbo_ said:
As for reducing the shutdown temp - I thought I had all that covered but I guess not! Once things are back up and running that's job 1.
My comment is assuming it was the heat that killed the drives before the temp reached the server's shutdown threshold.
A fair one i think, ambient temperature was 57 degrees - that's f***ing hot by anyones standard - lord alone knows what the internal temperatures were.
Jamie - WRT to cooling - there's a dedicated Aircon unit in that cupboard - but it packed in... Had I properly set the shutdown temperature I suspect I would have been saved a lot of heartache...
Anyway, 16GB of data later, and I've restored the mailserver to it's former state, total time taken to rebuild it was about 1hour 45minutes so I'm content with that.
As for changing capacity and raid setup - it's probably my next job come tomorrow!
JamieBeeston said:
![]()
GL With it.
Now.. to design a RAID Aircon Setup
Got one.
twin aircon linked by token ring I think. Each runs for 2 days then swaps to the other unless the load requires both. I one dies the other kicks in.
It was here when we moved in
servers are cool but the people are dropping in the heat as there's no aircon in the offices
>> Edited by malman on Monday 11th July 11:09
As a bit of a post mortem to my issue - here's the situation:
Dell diagnostics report two disks as "degraded". The actual raid controller reports no issues with the disks and a verify media runs with no errors.
So, clearly the dell utils see a problem that the raid controller can't see or fix. Can I recover these disks?
At the mo I'm thinking about bunging them into another machine with a different raid controller to see if I can get the data off them that way.
Dell diagnostics report two disks as "degraded". The actual raid controller reports no issues with the disks and a verify media runs with no errors.
So, clearly the dell utils see a problem that the raid controller can't see or fix. Can I recover these disks?
At the mo I'm thinking about bunging them into another machine with a different raid controller to see if I can get the data off them that way.
Gassing Station | Computers, Gadgets & Stuff | Top of Page | What's New | My Stuff


