And I'm changing this subject in case there are more replies, as the previous one kept causing me stress. Sent from my Mobile -----Original message----- From: Corby Schmitz <[email protected]> To: Dan Olson <[email protected]>, Ken Raffenetti <[email protected]> Cc: Hunter Matthews <[email protected]>, Core Admins <[email protected]>, Narayan Desai <[email protected]> Sent: Sat, Aug 13, 2011 00:22:14 GMT+00:00 Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL ** The rebuild just completed, so it's back to normal. We can deal with the online spare on Monday. Thanks again everyone for helping get this system put back together. Corby Schmitz Sent from my mobile office On Aug 12, 2011, at 13:42, Corby Schmitz <[email protected]> wrote:
I'm in 221 now doing some prep work.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:13, Corby Schmitz <[email protected]> wrote:
Much appreciated.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:02, Dan Olson <[email protected]> wrote:
OK.. Ken and I are heading over there to rack a server in a little bit anyway.. I'll dig up a drive and have it ready.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Dan Olson" <[email protected]> Cc: "Core Admins" <[email protected]>, "Narayan Desai" <[email protected]>, "Hunter Matthews" <[email protected]> Sent: Friday, August 12, 2011 1:00:14 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
221. I will be in around 2.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 12:59, Dan Olson <[email protected]> wrote:
I'm available. Where is this server?
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Narayan Desai" <[email protected]>, "Dan Olson" <[email protected]>, "Hunter Matthews" <[email protected]> Cc: "Core Admins" <[email protected]> Sent: Friday, August 12, 2011 12:31:29 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Anyone have cycles today?
- -corby
On Aug 11, 2011, at 11:27 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
If we have a drive, I will come in tomorrow and deal with this. It won't be till the afternoon as I am coming in on the red-eye and need to catch some sleep before surgery. If one of you will be around and are willing to give me a hand, that would be great.
Thanks all.
- -corby
On Aug 10, 2011, at 11:14 PM, Corby Schmitz wrote:
Sata.
Corby Schmitz Sent from my mobile office
On Aug 10, 2011, at 18:50, Hunter Matthews <[email protected]> wrote:
I can. I'll try and remember to see what I have in stock as well.
I assume these are 3.0Gb/s SATA drives? Or is that SAS?
-- Hunter Matthews Unix Administrator Office: Bldg 221 Room B240 Argonne National Labs, MCS Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 Never take candy from strangers. Especially on the internet.
On Aug 10, 2011, at 8:43 PM, Narayan Desai wrote:
I'd hope so. Can anyone verify? -nld
On Aug 10, 2011, at 4:48 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
the controller will just treat it like a 300, right?
- -corby
On Aug 10, 2011, at 2:47 PM, Narayan Desai wrote:
I don't have any this small. I could probably set you up with the 750 though. -nld
On Aug 10, 2011, at 9:52 AM, Dan Olson wrote:
If its a raid 10 you with no hot spare you should replace it ASAP. One more failure and the raid is gone. If it didn't get remounted the raid controller did its job and the os never saw the failure.
Its a good time to verify the backup logs also.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Core Admins" <[email protected]> Sent: Wednesday, August 10, 2011 9:37:28 AM Subject: Fwd: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Fluke just started throwing these errors last night. Looking at the system, the volume has not been remounted as RO. Does this mean that I have time and can the disk when I am back onsite next week. Anyone have thoughts on this?
cschmitz@fluke:~$ df -h Filesystem Size Used Avail Use% Mounted on /dev/sda1 19G 3.2G 15G 18% / varrun 1007M 168K 1007M 1% /var/run varlock 1007M 0 1007M 0% /var/lock udev 1007M 52K 1007M 1% /dev devshm 1007M 0 1007M 0% /dev/shm /dev/sdb1 512G 1.1G 486G 1% /disks /dev/sda7 50G 218M 48G 1% /sandbox /dev/sda5 3.7G 72M 3.5G 2% /tmp /dev/sda6 3.7G 678M 2.9G 19% /var/log
cschmitz@fluke:~$ mount /dev/sda1 on / type ext3 (rw,relatime,errors=remount-ro) proc on /proc type proc (rw,noexec,nosuid,nodev) /sys on /sys type sysfs (rw,noexec,nosuid,nodev) varrun on /var/run type tmpfs (rw,noexec,nosuid,nodev,mode=0755) varlock on /var/lock type tmpfs (rw,noexec,nosuid,nodev,mode=1777) udev on /dev type tmpfs (rw,mode=0755) devshm on /dev/shm type tmpfs (rw) devpts on /dev/pts type devpts (rw,gid=5,mode=620) /dev/sdb1 on /disks type ext3 (rw,relatime) /dev/sda7 on /sandbox type ext3 (rw,relatime) /dev/sda5 on /tmp type ext3 (rw,relatime) /dev/sda6 on /var/log type ext3 (rw,relatime) securityfs on /sys/kernel/security type securityfs (rw)
- -corby
Begin forwarded message:
From: [email protected] (Nagios on Fluke) Date: August 10, 2011 6:35:21 AM PDT To: [email protected] Subject: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
***** Nagios Running on fluke.anchor.anl.gov *****
Notification Type: PROBLEM
Service: 3WARE-RAID Host: localhost Address: 127.0.0.1 State: CRITICAL
Date/Time: Wed Aug 10 08:35:21 CDT 2011
Additional Info:
RAID CRITICAL: 1 array not OK - Array 0 status is DEGRADED(RAID-10 on adapter 8) URL: https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=l... calhost&service=3WARE-RAID
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFOQpeoQhpwH3ALVFERAh40AKCh6S42E0EZ96RbKhTzFNOOT3VWmwCeOiiG ecKSeHI0FPY2GANILo+Ej2U= =FuHI -----END PGP SIGNATURE-----
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFOQvyuQhpwH3ALVFERAuFRAKCOw9bIO6NFMpQdaQHNsbTk4FTv3QCeMjMl Fs3bMNGCBxSlbA5OE33JikI= =pFiR -----END PGP SIGNATURE-----
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFORKvDQhpwH3ALVFERApqKAJoCstK+D803FHQAV0PSVp61IGwFvACgwPKK /1JFOm914s/uWeP/W4Gb+tg= =g0S2 -----END PGP SIGNATURE-----
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFORWNxQhpwH3ALVFERAmrmAKCwDYTi48q7HZQx0Un/CWzTE1YWSwCgnNo5 if5IL0u/9YnXNTMfcMpMtXM= =TLgA -----END PGP SIGNATURE-----
I added the 5th drive as a spare. Should be all set. Unit UnitType Status %RCmpl %V/I/M Stripe Size(GB) Cache AVrfy ------------------------------------------------------------------------------ u0 RAID-10 OK - - 256K 596.025 RiW ON u1 SPARE OK - - - 698.629 - OFF VPort Status Unit Size Type Phy Encl-Slot Model ------------------------------------------------------------------------------ p0 OK u0 298.09 GB SATA 0 - WDC WD3200YS-01PGB0 p1 OK u0 298.09 GB SATA 1 - WDC WD3200YS-01PGB0 p2 OK u0 698.63 GB SATA 2 - WDC WD7500AYYS-01RC p3 OK u0 298.09 GB SATA 3 - WDC WD3200YS-01PGB0 p4 OK u1 698.63 GB SATA 4 - WDC WD7500AYYS-01RC Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 08/12/2011 07:38 PM, Craig Stacey wrote:
And I'm changing this subject in case there are more replies, as the previous one kept causing me stress.
/Sent from my Mobile/
-----Original message-----
*From: *Corby Schmitz <[email protected]>* To: *Dan Olson <[email protected]>, Ken Raffenetti <[email protected]>* Cc: *Hunter Matthews <[email protected]>, Core Admins <[email protected]>, Narayan Desai <[email protected]>* Sent: *Sat, Aug 13, 2011 00:22:14 GMT+00:00* Subject: *Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
The rebuild just completed, so it's back to normal. We can deal with the online spare on Monday. Thanks again everyone for helping get this system put back together.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:42, Corby Schmitz wrote:
> I'm in 221 now doing some prep work. > > Corby Schmitz > Sent from my mobile office > > On Aug 12, 2011, at 13:13, Corby Schmitz wrote: > >> Much appreciated. >> >> Corby Schmitz >> Sent from my mobile office >> >> On Aug 12, 2011, at 13:02, Dan Olson wrote: >> >>> OK.. Ken and I are heading over there to rack a server in a little bit anyway.. I'll dig up a drive and have it ready. >>> >>> ---- >>> Daniel Murphy-Olson >>> Systems Administrator >>> Mathematics & Computer Science Division >>> Argonne National Laboratory >>> 630-252-0055 >>> >>> ----- Original Message ----- >>> From: "Corby Schmitz" >>> To: "Dan Olson" >>> Cc: "Core Admins" , "Narayan Desai" , "Hunter Matthews" >>> Sent: Friday, August 12, 2011 1:00:14 PM >>> Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL ** >>> >>> 221. I will be in around 2. >>> >>> >>> Corby Schmitz >>> Sent from my mobile office >>> >>> On Aug 12, 2011, at 12:59, Dan Olson wrote: >>> >>>> I'm available. Where is this server? >>>> >>>> ---- >>>> Daniel Murphy-Olson >>>> Systems Administrator >>>> Mathematics & Computer Science Division >>>> Argonne National Laboratory >>>> 630-252-0055 >>>> >>>> ----- Original Message ----- >>>> From: "Schmitz Corby" >>>> To: "Narayan Desai" , "Dan Olson" , "Hunter Matthews" >>>> Cc: "Core Admins" >>>> Sent: Friday, August 12, 2011 12:31:29 PM >>>> Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL ** >>>> Anyone have cycles today?
-corby
On Aug 11, 2011, at 11:27 PM, Schmitz Corby wrote:
>>>>> -----BEGIN PGP SIGNED MESSAGE----- >>>>> Hash: SHA1 >>>>> >>>>> If we have a drive, I will come in tomorrow and deal with this. It won't be till the afternoon as I am coming in on the red-eye and need to catch some sleep before surgery. If one of you will be around and are willing to give me a hand, that would be great. >>>>> >>>>> Thanks all. >>>>> >>>>> - -corby >>>>> >>>>> >>>>> >>>>> On Aug 10, 2011, at 11:14 PM, Corby Schmitz wrote: >>>>> >>>>>> Sata. >>>>>> >>>>>> Corby Schmitz >>>>>> Sent from my mobile office >>>>>> >>>>>> On Aug 10, 2011, at 18:50, Hunter Matthews wrote: >>>>>> >>>>>>> I can. I'll try and remember to see what I have in stock as well. >>>>>>> >>>>>>> I assume these are 3.0Gb/s SATA drives? Or is that SAS? >>>>>>> >>>>>>> -- >>>>>>> Hunter Matthews Unix Administrator >>>>>>> Office: Bldg 221 Room B240 Argonne National Labs, MCS >>>>>>> Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 >>>>>>> Never take candy from strangers. Especially on the internet. >>>>>>> >>>>>>> On Aug 10, 2011, at 8:43 PM, Narayan Desai wrote: >>>>>>> >>>>>>>> I'd hope so. Can anyone verify? >>>>>>>> -nld >>>>>>>> >>>>>>>> On Aug 10, 2011, at 4:48 PM, Schmitz Corby wrote: >>>>>>>> >>>>>>>>> -----BEGIN PGP SIGNED MESSAGE----- >>>>>>>>> Hash: SHA1 >>>>>>>>> >>>>>>>>> the controller will just treat it like a 300, right? >>>>>>>>> >>>>>>>>> - -corby >>>>>>>>> >>>>>>>>> >>>>>>>>> >>>>>>>>> On Aug 10, 2011, at 2:47 PM, Narayan Desai wrote: >>>>>>>>> >>>>>>>>>> I don't have any this small. I could probably set you up with the 750 though. >>>>>>>>>> -nld >>>>>>>>>> >>>>>>>>>> On Aug 10, 2011, at 9:52 AM, Dan Olson wrote: >>>>>>>>>> >>>>>>>>>>> If its a raid 10 you with no hot spare you should replace it ASAP. One more failure and the raid is gone. If it didn't get remounted the raid controller did its job and the os never saw the failure. >>>>>>>>>>> >>>>>>>>>>> Its a good time to verify the backup logs also. >>>>>>>>>>> >>>>>>>>>>> ---- >>>>>>>>>>> Daniel Murphy-Olson >>>>>>>>>>> Systems Administrator >>>>>>>>>>> Mathematics & Computer Science Division >>>>>>>>>>> Argonne National Laboratory >>>>>>>>>>> 630-252-0055 >>>>>>>>>>> >>>>>>>>>>> ----- Original Message ----- >>>>>>>>>>> From: "Schmitz Corby" >>>>>>>>>>> To: "Core Admins" >>>>>>>>>>> Sent: Wednesday, August 10, 2011 9:37:28 AM >>>>>>>>>>> Subject: Fwd: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL ** >>>>>>>>>>> >>>>>>>>>>> -----BEGIN PGP SIGNED MESSAGE----- >>>>>>>>>>> Hash: SHA1 >>>>>>>>>>> >>>>>>>>>>> Fluke just started throwing these errors last night. Looking at the system, the volume has not been remounted as RO. Does this mean that I have time and can the disk when I am back onsite next week. Anyone have thoughts on this? >>>>>>>>>>> >>>>>>>>>>> cschmitz@fluke:~$ df -h >>>>>>>>>>> Filesystem Size Used Avail Use% Mounted on >>>>>>>>>>> /dev/sda1 19G 3.2G 15G 18% / >>>>>>>>>>> varrun 1007M 168K 1007M 1% /var/run >>>>>>>>>>> varlock 1007M 0 1007M 0% /var/lock >>>>>>>>>>> udev 1007M 52K 1007M 1% /dev >>>>>>>>>>> devshm 1007M 0 1007M 0% /dev/shm >>>>>>>>>>> /dev/sdb1 512G 1.1G 486G 1% /disks >>>>>>>>>>> /dev/sda7 50G 218M 48G 1% /sandbox >>>>>>>>>>> /dev/sda5 3.7G 72M 3.5G 2% /tmp >>>>>>>>>>> /dev/sda6 3.7G 678M 2.9G 19% /var/log >>>>>>>>>>> >>>>>>>>>>> cschmitz@fluke:~$ mount >>>>>>>>>>> /dev/sda1 on / type ext3 (rw,relatime,errors=remount-ro) >>>>>>>>>>> proc on /proc type proc (rw,noexec,nosuid,nodev) >>>>>>>>>>> /sys on /sys type sysfs (rw,noexec,nosuid,nodev) >>>>>>>>>>> varrun on /var/run type tmpfs (rw,noexec,nosuid,nodev,mode=0755) >>>>>>>>>>> varlock on /var/lock type tmpfs (rw,noexec,nosuid,nodev,mode=1777) >>>>>>>>>>> udev on /dev type tmpfs (rw,mode=0755) >>>>>>>>>>> devshm on /dev/shm type tmpfs (rw) >>>>>>>>>>> devpts on /dev/pts type devpts (rw,gid=5,mode=620) >>>>>>>>>>> /dev/sdb1 on /disks type ext3 (rw,relatime) >>>>>>>>>>> /dev/sda7 on /sandbox type ext3 (rw,relatime) >>>>>>>>>>> /dev/sda5 on /tmp type ext3 (rw,relatime) >>>>>>>>>>> /dev/sda6 on /var/log type ext3 (rw,relatime) >>>>>>>>>>> securityfs on /sys/kernel/security type securityfs (rw) >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> - -corby >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> Begin forwarded message: >>>>>>>>>>> >>>>>>>>>>>> From: [email protected] (Nagios on Fluke) >>>>>>>>>>>> Date: August 10, 2011 6:35:21 AM PDT >>>>>>>>>>>> To: [email protected] >>>>>>>>>>>> Subject: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL ** >>>>>>>>>>>> >>>>>>>>>>>> ***** Nagios Running on fluke.anchor.anl.gov ***** >>>>>>>>>>>> >>>>>>>>>>>> Notification Type: PROBLEM >>>>>>>>>>>> >>>>>>>>>>>> Service: 3WARE-RAID >>>>>>>>>>>> Host: localhost >>>>>>>>>>>> Address: 127.0.0.1 >>>>>>>>>>>> State: CRITICAL >>>>>>>>>>>> >>>>>>>>>>>> Date/Time: Wed Aug 10 08:35:21 CDT 2011 >>>>>>>>>>>> >>>>>>>>>>>> Additional Info: >>>>>>>>>>>> >>>>>>>>>>>> RAID CRITICAL: 1 array not OK - Array 0 status is DEGRADED(RAID-10 on adapter 8) >>>>>>>>>>>> URL: https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=l... <https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=localhost&service=3WARE-RAID> >>>>>>>>>>> >>>>>>>>>>>
>> >>
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 awesome. so it will be usable by either mirror as needed? - -corby On Aug 15, 2011, at 9:32 AM, Ken Raffenetti wrote:
I added the 5th drive as a spare. Should be all set.
Unit UnitType Status %RCmpl %V/I/M Stripe Size(GB) Cache AVrfy ------------------------------------------------------------------------------ u0 RAID-10 OK - - 256K 596.025 RiW ON u1 SPARE OK - - - 698.629 - OFF
VPort Status Unit Size Type Phy Encl-Slot Model ------------------------------------------------------------------------------ p0 OK u0 298.09 GB SATA 0 - WDC WD3200YS-01PGB0 p1 OK u0 298.09 GB SATA 1 - WDC WD3200YS-01PGB0 p2 OK u0 698.63 GB SATA 2 - WDC WD7500AYYS-01RC p3 OK u0 298.09 GB SATA 3 - WDC WD3200YS-01PGB0 p4 OK u1 698.63 GB SATA 4 - WDC WD7500AYYS-01RC
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 08/12/2011 07:38 PM, Craig Stacey wrote:
And I'm changing this subject in case there are more replies, as the previous one kept causing me stress.
/Sent from my Mobile/
-----Original message-----
*From: *Corby Schmitz <[email protected]>* To: *Dan Olson <[email protected]>, Ken Raffenetti <[email protected]>* Cc: *Hunter Matthews <[email protected]>, Core Admins <[email protected]>, Narayan Desai <[email protected]>* Sent: *Sat, Aug 13, 2011 00:22:14 GMT+00:00* Subject: *Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
The rebuild just completed, so it's back to normal. We can deal with the online spare on Monday. Thanks again everyone for helping get this system put back together.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:42, Corby Schmitz wrote:
I'm in 221 now doing some prep work.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:13, Corby Schmitz wrote:
Much appreciated.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:02, Dan Olson wrote:
OK.. Ken and I are heading over there to rack a server in a little bit anyway.. I'll dig up a drive and have it ready.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" To: "Dan Olson" Cc: "Core Admins" , "Narayan Desai" , "Hunter Matthews" Sent: Friday, August 12, 2011 1:00:14 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
221. I will be in around 2.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 12:59, Dan Olson wrote:
I'm available. Where is this server?
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" To: "Narayan Desai" , "Dan Olson" , "Hunter Matthews" Cc: "Core Admins" Sent: Friday, August 12, 2011 12:31:29 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
Anyone have cycles today?
-corby
On Aug 11, 2011, at 11:27 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
If we have a drive, I will come in tomorrow and deal with this. It won't be till the afternoon as I am coming in on the red-eye and need to catch some sleep before surgery. If one of you will be around and are willing to give me a hand, that would be great.
Thanks all.
- -corby
On Aug 10, 2011, at 11:14 PM, Corby Schmitz wrote:
Sata.
Corby Schmitz Sent from my mobile office
On Aug 10, 2011, at 18:50, Hunter Matthews wrote:
I can. I'll try and remember to see what I have in stock as well.
I assume these are 3.0Gb/s SATA drives? Or is that SAS?
-- Hunter Matthews Unix Administrator Office: Bldg 221 Room B240 Argonne National Labs, MCS Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 Never take candy from strangers. Especially on the internet.
On Aug 10, 2011, at 8:43 PM, Narayan Desai wrote:
I'd hope so. Can anyone verify? -nld
On Aug 10, 2011, at 4:48 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
the controller will just treat it like a 300, right?
- -corby
On Aug 10, 2011, at 2:47 PM, Narayan Desai wrote:
I don't have any this small. I could probably set you up with the 750 though. -nld
On Aug 10, 2011, at 9:52 AM, Dan Olson wrote:
If its a raid 10 you with no hot spare you should replace it ASAP. One more failure and the raid is gone. If it didn't get remounted the raid controller did its job and the os never saw the failure.
Its a good time to verify the backup logs also.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" To: "Core Admins" Sent: Wednesday, August 10, 2011 9:37:28 AM Subject: Fwd: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Fluke just started throwing these errors last night. Looking at the system, the volume has not been remounted as RO. Does this mean that I have time and can the disk when I am back onsite next week. Anyone have thoughts on this?
cschmitz@fluke:~$ df -h Filesystem Size Used Avail Use% Mounted on /dev/sda1 19G 3.2G 15G 18% / varrun 1007M 168K 1007M 1% /var/run varlock 1007M 0 1007M 0% /var/lock udev 1007M 52K 1007M 1% /dev devshm 1007M 0 1007M 0% /dev/shm /dev/sdb1 512G 1.1G 486G 1% /disks /dev/sda7 50G 218M 48G 1% /sandbox /dev/sda5 3.7G 72M 3.5G 2% /tmp /dev/sda6 3.7G 678M 2.9G 19% /var/log
cschmitz@fluke:~$ mount /dev/sda1 on / type ext3 (rw,relatime,errors=remount-ro) proc on /proc type proc (rw,noexec,nosuid,nodev) /sys on /sys type sysfs (rw,noexec,nosuid,nodev) varrun on /var/run type tmpfs (rw,noexec,nosuid,nodev,mode=0755) varlock on /var/lock type tmpfs (rw,noexec,nosuid,nodev,mode=1777) udev on /dev type tmpfs (rw,mode=0755) devshm on /dev/shm type tmpfs (rw) devpts on /dev/pts type devpts (rw,gid=5,mode=620) /dev/sdb1 on /disks type ext3 (rw,relatime) /dev/sda7 on /sandbox type ext3 (rw,relatime) /dev/sda5 on /tmp type ext3 (rw,relatime) /dev/sda6 on /var/log type ext3 (rw,relatime) securityfs on /sys/kernel/security type securityfs (rw)
- -corby
Begin forwarded message:
From: [email protected] (Nagios on Fluke) Date: August 10, 2011 6:35:21 AM PDT To: [email protected] Subject: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
***** Nagios Running on fluke.anchor.anl.gov *****
Notification Type: PROBLEM
Service: 3WARE-RAID Host: localhost Address: 127.0.0.1 State: CRITICAL
Date/Time: Wed Aug 10 08:35:21 CDT 2011
Additional Info:
RAID CRITICAL: 1 array not OK - Array 0 status is DEGRADED(RAID-10 on adapter 8) URL: https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=l... <https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=localhost&service=3WARE-RAID>
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFOSS6tQhpwH3ALVFERAm1vAJoDGoxMjBb6CWfQtnAxM1/HpnNkaACgvWP2 E9YAIwjch5EO2U/DJ8YFhU0= =rQ/I -----END PGP SIGNATURE-----
Supposedly. This is the way to do it according to all the docs I looked at. I haven't actually used one in production myself. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 08/15/2011 09:35 AM, Schmitz Corby wrote:
awesome. so it will be usable by either mirror as needed?
-corby
On Aug 15, 2011, at 9:32 AM, Ken Raffenetti wrote:
I added the 5th drive as a spare. Should be all set.
Unit UnitType Status %RCmpl %V/I/M Stripe Size(GB) Cache AVrfy ------------------------------------------------------------------------------ u0 RAID-10 OK - - 256K 596.025 RiW ON u1 SPARE OK - - - 698.629 - OFF
VPort Status Unit Size Type Phy Encl-Slot Model ------------------------------------------------------------------------------ p0 OK u0 298.09 GB SATA 0 - WDC WD3200YS-01PGB0 p1 OK u0 298.09 GB SATA 1 - WDC WD3200YS-01PGB0 p2 OK u0 698.63 GB SATA 2 - WDC WD7500AYYS-01RC p3 OK u0 298.09 GB SATA 3 - WDC WD3200YS-01PGB0 p4 OK u1 698.63 GB SATA 4 - WDC WD7500AYYS-01RC
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 08/12/2011 07:38 PM, Craig Stacey wrote:
And I'm changing this subject in case there are more replies, as the previous one kept causing me stress.
/Sent from my Mobile/
-----Original message-----
*From: *Corby Schmitz <[email protected]>* To: *Dan Olson <[email protected]>, Ken Raffenetti <[email protected]>* Cc: *Hunter Matthews <[email protected]>, Core Admins <[email protected]>, Narayan Desai <[email protected]>* Sent: *Sat, Aug 13, 2011 00:22:14 GMT+00:00* Subject: *Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
The rebuild just completed, so it's back to normal. We can deal with the online spare on Monday. Thanks again everyone for helping get this system put back together.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:42, Corby Schmitz wrote:
I'm in 221 now doing some prep work.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:13, Corby Schmitz wrote:
Much appreciated.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:02, Dan Olson wrote:
OK.. Ken and I are heading over there to rack a server in a little bit anyway.. I'll dig up a drive and have it ready.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" To: "Dan Olson" Cc: "Core Admins" , "Narayan Desai" , "Hunter Matthews" Sent: Friday, August 12, 2011 1:00:14 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
221. I will be in around 2.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 12:59, Dan Olson wrote:
I'm available. Where is this server?
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" To: "Narayan Desai" , "Dan Olson" , "Hunter Matthews" Cc: "Core Admins" Sent: Friday, August 12, 2011 12:31:29 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
Anyone have cycles today?
-corby
On Aug 11, 2011, at 11:27 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
If we have a drive, I will come in tomorrow and deal with this. It won't be till the afternoon as I am coming in on the red-eye and need to catch some sleep before surgery. If one of you will be around and are willing to give me a hand, that would be great.
Thanks all.
- -corby
On Aug 10, 2011, at 11:14 PM, Corby Schmitz wrote:
Sata.
Corby Schmitz Sent from my mobile office
On Aug 10, 2011, at 18:50, Hunter Matthews wrote:
I can. I'll try and remember to see what I have in stock as well.
I assume these are 3.0Gb/s SATA drives? Or is that SAS?
-- Hunter Matthews Unix Administrator Office: Bldg 221 Room B240 Argonne National Labs, MCS Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 Never take candy from strangers. Especially on the internet.
On Aug 10, 2011, at 8:43 PM, Narayan Desai wrote:
I'd hope so. Can anyone verify? -nld
On Aug 10, 2011, at 4:48 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
the controller will just treat it like a 300, right?
- -corby
On Aug 10, 2011, at 2:47 PM, Narayan Desai wrote:
I don't have any this small. I could probably set you up with the 750 though. -nld
On Aug 10, 2011, at 9:52 AM, Dan Olson wrote:
If its a raid 10 you with no hot spare you should replace it ASAP. One more failure and the raid is gone. If it didn't get remounted the raid controller did its job and the os never saw the failure.
Its a good time to verify the backup logs also.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" To: "Core Admins" Sent: Wednesday, August 10, 2011 9:37:28 AM Subject: Fwd: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Fluke just started throwing these errors last night. Looking at the system, the volume has not been remounted as RO. Does this mean that I have time and can the disk when I am back onsite next week. Anyone have thoughts on this?
cschmitz@fluke:~$ df -h Filesystem Size Used Avail Use% Mounted on /dev/sda1 19G 3.2G 15G 18% / varrun 1007M 168K 1007M 1% /var/run varlock 1007M 0 1007M 0% /var/lock udev 1007M 52K 1007M 1% /dev devshm 1007M 0 1007M 0% /dev/shm /dev/sdb1 512G 1.1G 486G 1% /disks /dev/sda7 50G 218M 48G 1% /sandbox /dev/sda5 3.7G 72M 3.5G 2% /tmp /dev/sda6 3.7G 678M 2.9G 19% /var/log
cschmitz@fluke:~$ mount /dev/sda1 on / type ext3 (rw,relatime,errors=remount-ro) proc on /proc type proc (rw,noexec,nosuid,nodev) /sys on /sys type sysfs (rw,noexec,nosuid,nodev) varrun on /var/run type tmpfs (rw,noexec,nosuid,nodev,mode=0755) varlock on /var/lock type tmpfs (rw,noexec,nosuid,nodev,mode=1777) udev on /dev type tmpfs (rw,mode=0755) devshm on /dev/shm type tmpfs (rw) devpts on /dev/pts type devpts (rw,gid=5,mode=620) /dev/sdb1 on /disks type ext3 (rw,relatime) /dev/sda7 on /sandbox type ext3 (rw,relatime) /dev/sda5 on /tmp type ext3 (rw,relatime) /dev/sda6 on /var/log type ext3 (rw,relatime) securityfs on /sys/kernel/security type securityfs (rw)
- -corby
Begin forwarded message:
From: [email protected] (Nagios on Fluke) Date: August 10, 2011 6:35:21 AM PDT To: [email protected] Subject: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
***** Nagios Running on fluke.anchor.anl.gov *****
Notification Type: PROBLEM
Service: 3WARE-RAID Host: localhost Address: 127.0.0.1 State: CRITICAL
Date/Time: Wed Aug 10 08:35:21 CDT 2011
Additional Info:
RAID CRITICAL: 1 array not OK - Array 0 status is DEGRADED(RAID-10 on adapter 8) URL: https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=l... <https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=localhost&service=3WARE-RAID>
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 sounds good to me. thanks. - -corby On Aug 15, 2011, at 9:36 AM, Ken Raffenetti wrote:
Supposedly. This is the way to do it according to all the docs I looked at. I haven't actually used one in production myself.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 08/15/2011 09:35 AM, Schmitz Corby wrote:
awesome. so it will be usable by either mirror as needed?
-corby
On Aug 15, 2011, at 9:32 AM, Ken Raffenetti wrote:
I added the 5th drive as a spare. Should be all set.
Unit UnitType Status %RCmpl %V/I/M Stripe Size(GB) Cache AVrfy ------------------------------------------------------------------------------ u0 RAID-10 OK - - 256K 596.025 RiW ON u1 SPARE OK - - - 698.629 - OFF
VPort Status Unit Size Type Phy Encl-Slot Model ------------------------------------------------------------------------------ p0 OK u0 298.09 GB SATA 0 - WDC WD3200YS-01PGB0 p1 OK u0 298.09 GB SATA 1 - WDC WD3200YS-01PGB0 p2 OK u0 698.63 GB SATA 2 - WDC WD7500AYYS-01RC p3 OK u0 298.09 GB SATA 3 - WDC WD3200YS-01PGB0 p4 OK u1 698.63 GB SATA 4 - WDC WD7500AYYS-01RC
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 08/12/2011 07:38 PM, Craig Stacey wrote:
And I'm changing this subject in case there are more replies, as the previous one kept causing me stress.
/Sent from my Mobile/
-----Original message-----
*From: *Corby Schmitz <[email protected]>* To: *Dan Olson <[email protected]>, Ken Raffenetti <[email protected]>* Cc: *Hunter Matthews <[email protected]>, Core Admins <[email protected]>, Narayan Desai <[email protected]>* Sent: *Sat, Aug 13, 2011 00:22:14 GMT+00:00* Subject: *Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
The rebuild just completed, so it's back to normal. We can deal with the online spare on Monday. Thanks again everyone for helping get this system put back together.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:42, Corby Schmitz wrote:
I'm in 221 now doing some prep work.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:13, Corby Schmitz wrote:
Much appreciated.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 13:02, Dan Olson wrote:
OK.. Ken and I are heading over there to rack a server in a little bit anyway.. I'll dig up a drive and have it ready.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" To: "Dan Olson" Cc: "Core Admins" , "Narayan Desai" , "Hunter Matthews" Sent: Friday, August 12, 2011 1:00:14 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
221. I will be in around 2.
Corby Schmitz Sent from my mobile office
On Aug 12, 2011, at 12:59, Dan Olson wrote:
I'm available. Where is this server?
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" To: "Narayan Desai" , "Dan Olson" , "Hunter Matthews" Cc: "Core Admins" Sent: Friday, August 12, 2011 12:31:29 PM Subject: Re: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
Anyone have cycles today?
-corby
On Aug 11, 2011, at 11:27 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
If we have a drive, I will come in tomorrow and deal with this. It won't be till the afternoon as I am coming in on the red-eye and need to catch some sleep before surgery. If one of you will be around and are willing to give me a hand, that would be great.
Thanks all.
- -corby
On Aug 10, 2011, at 11:14 PM, Corby Schmitz wrote:
Sata.
Corby Schmitz Sent from my mobile office
On Aug 10, 2011, at 18:50, Hunter Matthews wrote:
I can. I'll try and remember to see what I have in stock as well.
I assume these are 3.0Gb/s SATA drives? Or is that SAS?
-- Hunter Matthews Unix Administrator Office: Bldg 221 Room B240 Argonne National Labs, MCS Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 Never take candy from strangers. Especially on the internet.
On Aug 10, 2011, at 8:43 PM, Narayan Desai wrote:
I'd hope so. Can anyone verify? -nld
On Aug 10, 2011, at 4:48 PM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
the controller will just treat it like a 300, right?
- -corby
On Aug 10, 2011, at 2:47 PM, Narayan Desai wrote:
I don't have any this small. I could probably set you up with the 750 though. -nld
On Aug 10, 2011, at 9:52 AM, Dan Olson wrote:
If its a raid 10 you with no hot spare you should replace it ASAP. One more failure and the raid is gone. If it didn't get remounted the raid controller did its job and the os never saw the failure.
Its a good time to verify the backup logs also.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Schmitz Corby" To: "Core Admins" Sent: Wednesday, August 10, 2011 9:37:28 AM Subject: Fwd: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
Fluke just started throwing these errors last night. Looking at the system, the volume has not been remounted as RO. Does this mean that I have time and can the disk when I am back onsite next week. Anyone have thoughts on this?
cschmitz@fluke:~$ df -h Filesystem Size Used Avail Use% Mounted on /dev/sda1 19G 3.2G 15G 18% / varrun 1007M 168K 1007M 1% /var/run varlock 1007M 0 1007M 0% /var/lock udev 1007M 52K 1007M 1% /dev devshm 1007M 0 1007M 0% /dev/shm /dev/sdb1 512G 1.1G 486G 1% /disks /dev/sda7 50G 218M 48G 1% /sandbox /dev/sda5 3.7G 72M 3.5G 2% /tmp /dev/sda6 3.7G 678M 2.9G 19% /var/log
cschmitz@fluke:~$ mount /dev/sda1 on / type ext3 (rw,relatime,errors=remount-ro) proc on /proc type proc (rw,noexec,nosuid,nodev) /sys on /sys type sysfs (rw,noexec,nosuid,nodev) varrun on /var/run type tmpfs (rw,noexec,nosuid,nodev,mode=0755) varlock on /var/lock type tmpfs (rw,noexec,nosuid,nodev,mode=1777) udev on /dev type tmpfs (rw,mode=0755) devshm on /dev/shm type tmpfs (rw) devpts on /dev/pts type devpts (rw,gid=5,mode=620) /dev/sdb1 on /disks type ext3 (rw,relatime) /dev/sda7 on /sandbox type ext3 (rw,relatime) /dev/sda5 on /tmp type ext3 (rw,relatime) /dev/sda6 on /var/log type ext3 (rw,relatime) securityfs on /sys/kernel/security type securityfs (rw)
- -corby
Begin forwarded message:
From: [email protected] (Nagios on Fluke) Date: August 10, 2011 6:35:21 AM PDT To: [email protected] Subject: ** PROBLEM alert - localhost/3WARE-RAID is CRITICAL **
***** Nagios Running on fluke.anchor.anl.gov *****
Notification Type: PROBLEM
Service: 3WARE-RAID Host: localhost Address: 127.0.0.1 State: CRITICAL
Date/Time: Wed Aug 10 08:35:21 CDT 2011
Additional Info:
RAID CRITICAL: 1 array not OK - Array 0 status is DEGRADED(RAID-10 on adapter 8) URL: https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=l... <https://fluke.anchor.anl.gov/anchor/nagios/cgi-bin/extinfo.cgi?type=2&host=localhost&service=3WARE-RAID>
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFOSS/HQhpwH3ALVFERAhjvAKCKT/AEXC+NfaQ8kUn5o1+Uo4EJiQCfd5Gb K8jmiCI/+oXrw0VIWZ+mZ+c= =dv6P -----END PGP SIGNATURE-----
participants (3)
-
Craig Stacey -
Ken Raffenetti -
Schmitz Corby