List of failures from the cooling incident
I want to compile a list of failures that we feel were a direct result of the cooling loss. If you have hardware that failed as a result of this, or a significant loss of work, please let me know so I can pass this info up the chain. -- Craig
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 Fan tray in Fusion-force10 switch has a failed component. That is all on the network side. - -corby On Apr 19, 2012, at 11:09 AM, Craig Stacey wrote:
I want to compile a list of failures that we feel were a direct result of the cooling loss. If you have hardware that failed as a result of this, or a significant loss of work, please let me know so I can pass this info up the chain. -- Craig
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFPkDlhQhpwH3ALVFERApmMAJ9FyyVwWNOmPN4Yht6k5+qvOL5HHwCePQ/1 ibUarq2+/g2qZ0DU7lrvrNQ= =zkVo -----END PGP SIGNATURE-----
sto07 lost 2 fairly new drives (2tb hitachi 7k3000) ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Schmitz Corby" <[email protected]> To: "Craig Stacey" <[email protected]> Cc: "Core Admins" <[email protected]> Sent: Thursday, April 19, 2012 11:12:17 AM Subject: Re: List of failures from the cooling incident -----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 Fan tray in Fusion-force10 switch has a failed component. That is all on the network side. - -corby On Apr 19, 2012, at 11:09 AM, Craig Stacey wrote:
I want to compile a list of failures that we feel were a direct result of the cooling loss. If you have hardware that failed as a result of this, or a significant loss of work, please let me know so I can pass this info up the chain. -- Craig
-----BEGIN PGP SIGNATURE----- Version: GnuPG/MacGPG2 v2.0.16 (Darwin) iD8DBQFPkDlhQhpwH3ALVFERApmMAJ9FyyVwWNOmPN4Yht6k5+qvOL5HHwCePQ/1 ibUarq2+/g2qZ0DU7lrvrNQ= =zkVo -----END PGP SIGNATURE-----
On Thu, Apr 19, 2012 at 11:09:03AM -0500, Craig Stacey wrote:
I want to compile a list of failures that we feel were a direct result of the cooling loss. If you have hardware that failed as a result of this, or a significant loss of work, please let me know so I can pass this info up the chain.
We lost one power supply on a fusion iDataPlex compute node enclosure (one supply powers two nodes). It went out before fusion was shutdown, so it went out because of the heat, not because of a power off-power on cycle. And as Corby mentioned, we lost something (likely a fan) in the fan tray in fusion's Force10 switch. As far as lost work, fusion, kbt and cosmea were all full running jobs which had to be interrupted and restarted due to the shutdown. I don't know which jobs were able to restart from where they left off and so can't give an exact figure for amount of lost work, but rough estimate off the top of my head based on elapsed job walltime and averaged over the clusters was we lost approx 2 days (48 hours) of compute time. Then I lost my Sunday recovering from the outage, which I would have rather spend in other ways... :) John
It may be also worth noting that damage due to overheating isn't always immediate; it can show up days or months later as the result of premature failure of components as a result of overheating. Just right now, a fusion compute node uncharacteristically powered itself off (haven't had a node do that in the last 2 years of the cluster). Whether that was due to the heat, it's hard to say, but that suspicion will always be in the back of my mind from now on. John
The 10Gb uplink blade of PADS IB switch Failed HD in PADS admin node Failed HD in FG compute node That's it so far. On Apr 19, 2012, at 11:09 AM, Craig Stacey wrote:
I want to compile a list of failures that we feel were a direct result of the cooling loss. If you have hardware that failed as a result of this, or a significant loss of work, please let me know so I can pass this info up the chain. -- Craig
IGSB: 3 failed HD's, 1 compute node power supply, and about 2 days of jobs + developer time to put it back together. On Apr 19, 2012, at 12:59 PM, Ti Leggett wrote:
The 10Gb uplink blade of PADS IB switch Failed HD in PADS admin node Failed HD in FG compute node
That's it so far.
On Apr 19, 2012, at 11:09 AM, Craig Stacey wrote:
I want to compile a list of failures that we feel were a direct result of the cooling loss. If you have hardware that failed as a result of this, or a significant loss of work, please let me know so I can pass this info up the chain. -- Craig
-- Hunter Matthews Unix Administrator Office: Bldg 221 Room B240 Argonne National Labs, MCS Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 Never take candy from strangers. Especially on the internet.
We lost an SFP from the myricom switch, and we now have a module from the sicortex that is unusable because one of the temperature sensors constantly reports a 180C temp. This means that we are now completely out of spares for the sicortex, with no potential for replenishment. -nld On Apr 19, 2012, at 1:56 PM, Hunter Matthews wrote:
IGSB: 3 failed HD's, 1 compute node power supply, and about 2 days of jobs + developer time to put it back together.
On Apr 19, 2012, at 12:59 PM, Ti Leggett wrote:
The 10Gb uplink blade of PADS IB switch Failed HD in PADS admin node Failed HD in FG compute node
That's it so far.
On Apr 19, 2012, at 11:09 AM, Craig Stacey wrote:
I want to compile a list of failures that we feel were a direct result of the cooling loss. If you have hardware that failed as a result of this, or a significant loss of work, please let me know so I can pass this info up the chain. -- Craig
-- Hunter Matthews Unix Administrator Office: Bldg 221 Room B240 Argonne National Labs, MCS Key: F0F88438 / FFB5 34C0 B350 99A4 BB02 9779 A5DB 8B09 F0F8 8438 Never take candy from strangers. Especially on the internet.
participants (7)
-
Craig Stacey -
Dan Olson -
Hunter Matthews -
John Valdes -
Narayan Desai -
Schmitz Corby -
Ti Leggett