We got a bunch of software raid failures on Nagasaki. I'll deal with these when I get back. Sent from my iPhone On Apr 15, 2012, at 1:51 PM, John Valdes <[email protected]> wrote:
No problem; I set it to vibrate and dropped it in the martini mixer to stir my martinis. :)
John
On Sun, Apr 15, 2012 at 09:17:31AM -0500, Corby Schmitz wrote:
Sorry about that. It escaped me that acking alerts once everyone was aware of the problem was a good idea.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:44, John Valdes <[email protected]> wrote:
I'll be in tomorrow (afternoon) bringing fusion, kbt and cosmea back up (their DDNs and fileservers can't be powered back on remotely, unfortunately) and can poke other systems if needed too.
John
PS. We managed to fill the memory in the pager!
On Sun, Apr 15, 2012 at 12:36:25AM -0500, Craig Stacey wrote:
If you can get Hunter's, he'll love you to pieces.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]>, "Corby Schmitz" <[email protected]> Sent: Sunday, April 15, 2012 12:30:03 AM Subject: Re: Temp event in the core
The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core
I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core
Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time.
I'm heading to bed, but call the cell if you need something.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----