No objections from me. -- Craig ________________________________ From: Schmitz, Corby B. Sent: Thursday, February 4, 2016 7:43 PM To: Stacey, Craig Cc: helpinfo; CELS Admins ([email protected]); Sidorowicz, Kenneth V.; Leibfritz, David W.; [email protected]; [email protected]; Schmidt, Jack C.; [email protected]; Dust, Jim Subject: Re: [CELS Sysadmins] Network Outage Overnight - Card swap required No, just mcs proper. Unless you count their desktop net. Corby Schmitz Sent from my mobile office On Feb 4, 2016, at 7:24 PM, Stacey, Craig <[email protected]<mailto:[email protected]>> wrote: Does this affect LCF? -- Craig (mobile) From: Schmitz, Corby B. Sent: Feb 4, 2016 7:11 PM To: helpinfo; CELS Admins ([email protected]<mailto:[email protected]>); Sidorowicz, Kenneth V.; Leibfritz, David W.; [email protected]<mailto:[email protected]> Cc: [email protected]<mailto:[email protected]>; Schmidt, Jack C.; [email protected]<mailto:[email protected]>; Dust, Jim Subject: Re: [CELS Sysadmins] Network Outage Overnight - Card swap required All: I finally have resolution on this particular issue and I wanted to share what we know, and what we are doing with this knowledge. JTAC took a good deal of time working through the logs and debugging data. It was identified that the memory on the MPC in slot 3 in the border router is experiencing errors. Many of these are correctable. In the instance noted below, it was unable to correct the hardware issue. JTAC ran through other cases where this has happened and decided that the best course of action was to replace the card and move forward. We spent a ittle time ensuring that the errors were not coming from the fabric cards, but that turned out not to be the case. I have taken delivery of the replacement card under RMA and have submitted a change request (CHG#31762) for 0530hrs CST on Tuesday 2/9/2016. It is going through the final approval process within the CAB, but I wanted to check in, in parallel, to ensure that there were no concerns with swapping out the card at that time. The affected connections are listed below: Xe-3/0/0.803,804 – MREN Backup Link (R&E and TR/CPS) – Will not be visible Xe-3/0/0.345,346 – ESnet Tertiary Link – Will not be visible Xe-3/0/1.675 – MCS-240rtr – May experience a momentary blip of under 5 seconds Xe-3/2/3 – CyberSpan - Will lose replicated traffic during the maintenance window (expected to be less than 10 minutes) Xe-3/3/0.1300 – APS-ScienceDMZ – Will lose connectivity during the maintenance window (expected to be less than 10 minutes) Please let me know if there are any issues or concerns. If there is a more suitable timeframe for the swap, please suggest. corby ________________________________ From: Schmitz, Corby B. Sent: Wednesday, January 27, 2016 7:56 AM To: helpinfo; CELS Admins ([email protected]<mailto:[email protected]>); Sidorowicz, Kenneth V. ([email protected]<mailto:[email protected]>); Dave Leibfritz; [email protected]<mailto:[email protected]> Cc: [email protected]<mailto:[email protected]>; [email protected]<mailto:[email protected]>; Schmidt, Jack C.; Dust, Jim Subject: Network Outage Overnight All: At 0050hrs last night, the border router in 541b (noni) suffered a hardware-related event. Linecard #3 clocked a number of errors on traffic received from the chassis backplane. At 0057.16hrs, the chassis health monitor set the system status to Yellow, noting an impact to some traffic flows through linecard #3. At 0057.23hrs, the chassis health monitor set the system status to Red, indicating that the errors were unrecoverable in the current situation and that all traffic through linecard #3 was impacted. This culminated in a chassis health process initiating a module reset at 0057.38hrs, which ceased all traffic across linecard #3. At 0100.11hrs, linecard #3 had restored to operational state and by 0103.05hrs, all related network services had been fully restored. The following links were impacted by this outage: Xe-3/0/0.803,804 – MREN Backup Link (R&E and TR/CPS) – not active at the time of the event Xe-3/0/0.345,346 – ESnet Tertiary Link – not active at the time of the event Xe-3/0/1.675 – MCS-240rtr – networking on alternate path via lulo was not impacted, short outage of 15s was likely perceived Xe-3/0/2.668 – T1-Firewall link deprecated during maintenance weekend – not in production state Xe-3/1/1 – 308Core1-L2 – Layer 2 service backup link, all primary services were unaffected on 100G link to Core541 Xe-3/2/3 – CyberSpan, during this window the cyber monitoring on noni was non-functional – 0050-0103hrs Xe-3/3/0.1300 – APS-ScienceDMZ – impacted throughout the window from 0050-0103hrs We are working with Juniper TAC to identify the root cause of the hardware failure, and take whatever action needed to rectify the situation. There have been no additional indications of fabric errors since the card reset this morning, and no indication prior to the log messages at 0500hrs this morning. More information will be provided as we move through the TAC process. If you have any questions or concerns, please reach out to me directly, or send a note to [email protected]<mailto:[email protected]>. -- Corby Schmitz Network Communication Operations and Support Manager Computer and Information Systems Division - Network Engineer MREN/Starlight -- Argonne National Laboratory 9700 S. Cass Ave. Argonne, IL 60439 Desk: 630-252-7664 Cell: 630-296-4252 E-mail: [email protected]<mailto:[email protected]>