I think we have a temp event in the core. Don't know who to call. Corby
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile) Corby Schmitz <[email protected]> wrote: I think we have a temp event in the core. Don't know who to call. Corby
Cool. Wasn't sure if we contacted building mgmt first or not. Thanks. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:39, Craig Stacey <[email protected]> wrote: I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile) Corby Schmitz <[email protected]> wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
I don't even have their numbers. -- Craig ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core Cool. Wasn't sure if we contacted building mgmt first or not. Thanks. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote: I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile) Corby Schmitz < [email protected] > wrote: I think we have a temp event in the core. Don't know who to call. Corby
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php Looks like it started climbing around 8:30p. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
I'm starting to shut down things. First up is tape drives / halting tsm. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php Looks like it started climbing around 8:30p. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
If you start dropping VMs and such, try to leave watcher and net man running. It will make it easier to keep an eye on status and keep the network running. If non, just give me a heads up. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:55, Dan Olson <[email protected]> wrote:
I'm starting to shut down things. First up is tape drives / halting tsm.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php
Looks like it started climbing around 8:30p.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
For those following along at home Beagle already shut itself down. Ti's getting PADS and FG (and the DDN). Hopefully, Rene has raised the right people by now. I'm disheartened that *we* are still discovering cooling failures sooner than the building engineers. -- Craig ----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Craig Stacey" <[email protected]> Sent: Saturday, April 14, 2012 9:54:59 PM Subject: Re: Temp event in the core I'm starting to shut down things. First up is tape drives / halting tsm. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php Looks like it started climbing around 8:30p. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
Please tell me it was email that alerted you and not my message to the pager. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:58, Craig Stacey <[email protected]> wrote:
For those following along at home Beagle already shut itself down. Ti's getting PADS and FG (and the DDN).
Hopefully, Rene has raised the right people by now.
I'm disheartened that *we* are still discovering cooling failures sooner than the building engineers.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Craig Stacey" <[email protected]> Sent: Saturday, April 14, 2012 9:54:59 PM Subject: Re: Temp event in the core
I'm starting to shut down things. First up is tape drives / halting tsm.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php
Looks like it started climbing around 8:30p.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
My notification was your mail at the bottom of this thread. -- Craig ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:00:33 PM Subject: Re: Temp event in the core Please tell me it was email that alerted you and not my message to the pager. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 21:58, Craig Stacey <[email protected]> wrote:
For those following along at home Beagle already shut itself down. Ti's getting PADS and FG (and the DDN).
Hopefully, Rene has raised the right people by now.
I'm disheartened that *we* are still discovering cooling failures sooner than the building engineers.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Craig Stacey" <[email protected]> Sent: Saturday, April 14, 2012 9:54:59 PM Subject: Re: Temp event in the core
I'm starting to shut down things. First up is tape drives / halting tsm.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php
Looks like it started climbing around 8:30p.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti I am confused as to why these messages didn't go out. I see them in mail.info on watcher. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 22:04, Craig Stacey <[email protected]> wrote:
My notification was your mail at the bottom of this thread.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:00:33 PM Subject: Re: Temp event in the core
Please tell me it was email that alerted you and not my message to the pager.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:58, Craig Stacey <[email protected]> wrote:
For those following along at home Beagle already shut itself down. Ti's getting PADS and FG (and the DDN).
Hopefully, Rene has raised the right people by now.
I'm disheartened that *we* are still discovering cooling failures sooner than the building engineers.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Craig Stacey" <[email protected]> Sent: Saturday, April 14, 2012 9:54:59 PM Subject: Re: Temp event in the core
I'm starting to shut down things. First up is tape drives / halting tsm.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php
Looks like it started climbing around 8:30p.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
I've shut down ~350 magellan systems. I'll keep my eye on this thread and start powering down active systems if necessary. Jason On Apr 14, 2012, at 10:13 PM, Corby Schmitz wrote:
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti
I am confused as to why these messages didn't go out. I see them in mail.info on watcher.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 22:04, Craig Stacey <[email protected]> wrote:
My notification was your mail at the bottom of this thread.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:00:33 PM Subject: Re: Temp event in the core
Please tell me it was email that alerted you and not my message to the pager.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:58, Craig Stacey <[email protected]> wrote:
For those following along at home Beagle already shut itself down. Ti's getting PADS and FG (and the DDN).
Hopefully, Rene has raised the right people by now.
I'm disheartened that *we* are still discovering cooling failures sooner than the building engineers.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Craig Stacey" <[email protected]> Sent: Saturday, April 14, 2012 9:54:59 PM Subject: Re: Temp event in the core
I'm starting to shut down things. First up is tape drives / halting tsm.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php
Looks like it started climbing around 8:30p.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
To what address did it send for me? -- Craig ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:13:19 PM Subject: Re: Temp event in the core I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti I am confused as to why these messages didn't go out. I see them in mail.info on watcher. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 22:04, Craig Stacey <[email protected]> wrote:
My notification was your mail at the bottom of this thread.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:00:33 PM Subject: Re: Temp event in the core
Please tell me it was email that alerted you and not my message to the pager.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:58, Craig Stacey <[email protected]> wrote:
For those following along at home Beagle already shut itself down. Ti's getting PADS and FG (and the DDN).
Hopefully, Rene has raised the right people by now.
I'm disheartened that *we* are still discovering cooling failures sooner than the building engineers.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Craig Stacey" <[email protected]> Sent: Saturday, April 14, 2012 9:54:59 PM Subject: Re: Temp event in the core
I'm starting to shut down things. First up is tape drives / halting tsm.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php
Looks like it started climbing around 8:30p.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
[email protected] We have started to cross critical thresholds, so it has begun to send to pagers, [email protected] My phone is blowing up now. Corby Schmitz Sent from my mobile office On Apr 14, 2012, at 22:26, Craig Stacey <[email protected]> wrote:
To what address did it send for me?
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:13:19 PM Subject: Re: Temp event in the core
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti
I am confused as to why these messages didn't go out. I see them in mail.info on watcher.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 22:04, Craig Stacey <[email protected]> wrote:
My notification was your mail at the bottom of this thread.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:00:33 PM Subject: Re: Temp event in the core
Please tell me it was email that alerted you and not my message to the pager.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:58, Craig Stacey <[email protected]> wrote:
For those following along at home Beagle already shut itself down. Ti's getting PADS and FG (and the DDN).
Hopefully, Rene has raised the right people by now.
I'm disheartened that *we* are still discovering cooling failures sooner than the building engineers.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Craig Stacey" <[email protected]> Sent: Saturday, April 14, 2012 9:54:59 PM Subject: Re: Temp event in the core
I'm starting to shut down things. First up is tape drives / halting tsm.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: [email protected] Sent: Saturday, April 14, 2012 9:51:43 PM Subject: Re: Temp event in the core
Nice. I just looked at the temp map, and it isn't pretty: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php
Looks like it started climbing around 8:30p.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:43, Craig Stacey <[email protected]> wrote:
I don't even have their numbers.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Corby Schmitz" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 9:42:28 PM Subject: Re: Temp event in the core
Cool. Wasn't sure if we contacted building mgmt first or not.
Thanks.
Corby Schmitz Sent from my mobile office
On Apr 14, 2012, at 21:39, Craig Stacey < [email protected] > wrote:
I'll call Rene whose number is on the systems wiki. -- Craig (from my mobile)
Corby Schmitz < [email protected] > wrote:
I think we have a temp event in the core. Don't know who to call.
Corby
On Sat, Apr 14, 2012 at 10:13:19PM -0500, Corby Schmitz wrote:
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti
I am confused as to why these messages didn't go out. I see them in mail.info on watcher.
FYI, I'm receiving the watcher/nagios notices in email, and MCS pager is receiving notices as well, although the first notice to the pager was your email. John
This was my fault. The nagios alerts were getting caught in a filter. They're not now (at least, not temp/power related ones), as evidenced by my frantic acknowledging. My phone has not yet received any pages, though, which is supposed to happen, but we can troubleshoot that on another day. -- Craig ----- Original Message ----- From: "John Valdes" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:59:51 PM Subject: Re: Temp event in the core On Sat, Apr 14, 2012 at 10:13:19PM -0500, Corby Schmitz wrote:
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti
I am confused as to why these messages didn't go out. I see them in mail.info on watcher.
FYI, I'm receiving the watcher/nagios notices in email, and MCS pager is receiving notices as well, although the first notice to the pager was your email. John
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 sure. lets work on it on monday. - -corby On Apr 14, 2012, at 11:03 PM, Craig Stacey wrote:
This was my fault. The nagios alerts were getting caught in a filter. They're not now (at least, not temp/power related ones), as evidenced by my frantic acknowledging.
My phone has not yet received any pages, though, which is supposed to happen, but we can troubleshoot that on another day.
-- Craig
----- Original Message ----- From: "John Valdes" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:59:51 PM Subject: Re: Temp event in the core
On Sat, Apr 14, 2012 at 10:13:19PM -0500, Corby Schmitz wrote:
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti
I am confused as to why these messages didn't go out. I see them in mail.info on watcher.
FYI, I'm receiving the watcher/nagios notices in email, and MCS pager is receiving notices as well, although the first notice to the pager was your email.
John
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFPikj8QhpwH3ALVFERAtXZAKCx68LV4sTjRdZJEFx5Hilly/u5BACff2J6 Sb5+ZKNavYDvviwc3keOD+s= =vj5t -----END PGP SIGNATURE-----
I think all of IGSB is down, with a few exceptions noted below: I left kasserine-5, waterloo, leipzig and austerlitz running on purpose - the Gene Sequencer (HiSeq) is running and I would very much prefer not to mess with that - it gets expensive if we loose a run. In the case of direst emergency, I think we can shutdown leipzig and kasserine-5: The sequencer is writing to austerlitz and just needs waterloo for network stuff ( and waterloo is a cool runner, so wont add much heat). If someone could turn off the fiber channel storage array that the tokyo-1 and -2 machines are connected to that will help - it doesn't have a remote power off function. Normandy-1 and its 4 disk units are similar - I can't really power it off remotely. I think at least a couple of our machines powered off on their own due to heat. On Apr 15, 2012, at 12:05 AM, Schmitz Corby wrote:
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
sure. lets work on it on monday.
- -corby
On Apr 14, 2012, at 11:03 PM, Craig Stacey wrote:
This was my fault. The nagios alerts were getting caught in a filter. They're not now (at least, not temp/power related ones), as evidenced by my frantic acknowledging.
My phone has not yet received any pages, though, which is supposed to happen, but we can troubleshoot that on another day.
-- Craig
----- Original Message ----- From: "John Valdes" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: [email protected], "Dan Olson" <[email protected]>, [email protected] Sent: Saturday, April 14, 2012 10:59:51 PM Subject: Re: Temp event in the core
On Sat, Apr 14, 2012 at 10:13:19PM -0500, Corby Schmitz wrote:
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti
I am confused as to why these messages didn't go out. I see them in mail.info on watcher.
FYI, I'm receiving the watcher/nagios notices in email, and MCS pager is receiving notices as well, although the first notice to the pager was your email.
John
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPikj8QhpwH3ALVFERAtXZAKCx68LV4sTjRdZJEFx5Hilly/u5BACff2J6 Sb5+ZKNavYDvviwc3keOD+s= =vj5t -----END PGP SIGNATURE-----
-- Hunter Matthews Unix Administrator Office: Bldg 221 Room B240 Argonne National Labs, MCS Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 Never take candy from strangers. Especially on the internet.
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why. - -corby -----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
Seeing many recovery messages... -- Craig ----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core -----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why. - -corby -----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time. I'm heading to bed, but call the cell if you need something. Corby Schmitz Sent from my mobile office On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air. -- Craig ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time. I'm heading to bed, but call the cell if you need something. Corby Schmitz Sent from my mobile office On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air. -- Craig ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time. I'm heading to bed, but call the cell if you need something. Corby Schmitz Sent from my mobile office On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
If you can get Hunter's, he'll love you to pieces. -- Craig ----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]>, "Corby Schmitz" <[email protected]> Sent: Sunday, April 15, 2012 12:30:03 AM Subject: Re: Temp event in the core The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air. -- Craig ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time. I'm heading to bed, but call the cell if you need something. Corby Schmitz Sent from my mobile office On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
I'll be in tomorrow (afternoon) bringing fusion, kbt and cosmea back up (their DDNs and fileservers can't be powered back on remotely, unfortunately) and can poke other systems if needed too. John PS. We managed to fill the memory in the pager! On Sun, Apr 15, 2012 at 12:36:25AM -0500, Craig Stacey wrote:
If you can get Hunter's, he'll love you to pieces.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]>, "Corby Schmitz" <[email protected]> Sent: Sunday, April 15, 2012 12:30:03 AM Subject: Re: Temp event in the core
The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core
I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core
Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time.
I'm heading to bed, but call the cell if you need something.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
Sorry about that. It escaped me that acking alerts once everyone was aware of the problem was a good idea. Corby Schmitz Sent from my mobile office On Apr 15, 2012, at 0:44, John Valdes <[email protected]> wrote:
I'll be in tomorrow (afternoon) bringing fusion, kbt and cosmea back up (their DDNs and fileservers can't be powered back on remotely, unfortunately) and can poke other systems if needed too.
John
PS. We managed to fill the memory in the pager!
On Sun, Apr 15, 2012 at 12:36:25AM -0500, Craig Stacey wrote:
If you can get Hunter's, he'll love you to pieces.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]>, "Corby Schmitz" <[email protected]> Sent: Sunday, April 15, 2012 12:30:03 AM Subject: Re: Temp event in the core
The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core
I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core
Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time.
I'm heading to bed, but call the cell if you need something.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
I'm in the datacenter right now powering up MCS and SEED. I'm going to look at mg-rast next. Let me know if there is something you want powered on. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: "John Valdes" <[email protected]> Cc: "Craig Stacey" <[email protected]>, "Dan Olson" <[email protected]>, "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 10:17:31 AM Subject: Re: Temp event in the core Sorry about that. It escaped me that acking alerts once everyone was aware of the problem was a good idea. Corby Schmitz Sent from my mobile office On Apr 15, 2012, at 0:44, John Valdes <[email protected]> wrote:
I'll be in tomorrow (afternoon) bringing fusion, kbt and cosmea back up (their DDNs and fileservers can't be powered back on remotely, unfortunately) and can poke other systems if needed too.
John
PS. We managed to fill the memory in the pager!
On Sun, Apr 15, 2012 at 12:36:25AM -0500, Craig Stacey wrote:
If you can get Hunter's, he'll love you to pieces.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]>, "Corby Schmitz" <[email protected]> Sent: Sunday, April 15, 2012 12:30:03 AM Subject: Re: Temp event in the core
The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core
I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core
Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time.
I'm heading to bed, but call the cell if you need something.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
No problem; I set it to vibrate and dropped it in the martini mixer to stir my martinis. :) John On Sun, Apr 15, 2012 at 09:17:31AM -0500, Corby Schmitz wrote:
Sorry about that. It escaped me that acking alerts once everyone was aware of the problem was a good idea.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:44, John Valdes <[email protected]> wrote:
I'll be in tomorrow (afternoon) bringing fusion, kbt and cosmea back up (their DDNs and fileservers can't be powered back on remotely, unfortunately) and can poke other systems if needed too.
John
PS. We managed to fill the memory in the pager!
On Sun, Apr 15, 2012 at 12:36:25AM -0500, Craig Stacey wrote:
If you can get Hunter's, he'll love you to pieces.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]>, "Corby Schmitz" <[email protected]> Sent: Sunday, April 15, 2012 12:30:03 AM Subject: Re: Temp event in the core
The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core
I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core
Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time.
I'm heading to bed, but call the cell if you need something.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
We got a bunch of software raid failures on Nagasaki. I'll deal with these when I get back. Sent from my iPhone On Apr 15, 2012, at 1:51 PM, John Valdes <[email protected]> wrote:
No problem; I set it to vibrate and dropped it in the martini mixer to stir my martinis. :)
John
On Sun, Apr 15, 2012 at 09:17:31AM -0500, Corby Schmitz wrote:
Sorry about that. It escaped me that acking alerts once everyone was aware of the problem was a good idea.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:44, John Valdes <[email protected]> wrote:
I'll be in tomorrow (afternoon) bringing fusion, kbt and cosmea back up (their DDNs and fileservers can't be powered back on remotely, unfortunately) and can poke other systems if needed too.
John
PS. We managed to fill the memory in the pager!
On Sun, Apr 15, 2012 at 12:36:25AM -0500, Craig Stacey wrote:
If you can get Hunter's, he'll love you to pieces.
-- Craig
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]>, "Corby Schmitz" <[email protected]> Sent: Sunday, April 15, 2012 12:30:03 AM Subject: Re: Temp event in the core
The MCS / seed systems that I powered down need to be poked to come back up, I'm planning on heading in at 8am or so.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Corby Schmitz" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:45 AM Subject: Re: Temp event in the core
I'm suggesting we keep stuff off until tomorrow. The temps make me think they got outside air coming in, but not chilled air.
-- Craig
----- Original Message ----- From: "Corby Schmitz" <[email protected]> To: [email protected] Cc: "core-admins" <[email protected]> Sent: Sunday, April 15, 2012 12:17:07 AM Subject: Re: Temp event in the core
Graphs and maps look better. The routers just dropped below the critical threshold, so it was just in time.
I'm heading to bed, but call the cell if you need something.
Corby Schmitz Sent from my mobile office
On Apr 15, 2012, at 0:03, Craig Stacey <[email protected]> wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
Looks like Chiller2 has come back, judging from the temp map: https://watcher.mcs.anl.gov/mcs/temps/ssf-zone1.php John On Sun, Apr 15, 2012 at 12:03:40AM -0500, Craig Stacey wrote:
Seeing many recovery messages...
-- Craig
----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "core-admins" <[email protected]> Sent: Saturday, April 14, 2012 11:48:09 PM Subject: Re: Temp event in the core
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1
We just crossed into a new area of trouble. The 10GigE cards in the HPC and MCS routers have reached critical level 1 which means they are now operating at 65C or higher. At 75C they shut down. At that point, we would lose all connectivity outside connectivity. I would have to go in and reload the boxes to restore when the temp comes down. There is nothing we can do as these boxes cannot be shut down remotely, just rebooted. Just a heads up. If it happens, we know why.
- -corby
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin)
iD8DBQFPilMJQhpwH3ALVFERAtEyAJ4v07hN1osj//Sz7wM8PO76Z2SxBwCgkdR2 gmCq1DNC7/6Wlfp/Nidx9ww= =B3ve -----END PGP SIGNATURE-----
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 as it turns out we don't escalate until it hits a critical level. its email only on warnings. I jumped the gun on my pager notice I guess. - -corby On Apr 14, 2012, at 10:59 PM, John Valdes wrote:
On Sat, Apr 14, 2012 at 10:13:19PM -0500, Corby Schmitz wrote:
I'm looking at watchers notify config and it lists: Corby,max,Craig,Dan,John,ti
I am confused as to why these messages didn't go out. I see them in mail.info on watcher.
FYI, I'm receiving the watcher/nagios notices in email, and MCS pager is receiving notices as well, although the first notice to the pager was your email.
John
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFPikiDQhpwH3ALVFERAo3VAKDVnd8BByZtTfdT9FLfePcuflg1iwCcCngw +don1Oq7V1sBGw07ZsJyHr0= =jJCY -----END PGP SIGNATURE-----
participants (8)
-
Corby Schmitz -
Corby Schmitz -
Craig Stacey -
Dan Olson -
Hunter Matthews -
Jason Hedden -
John Valdes -
Schmitz Corby