Xe-3/0/0.803,804 – MREN Backup Link (R&E and TR/CPS) – Will not be visible
Xe-3/0/0.345,346 – ESnet Tertiary Link – Will not be visible
Xe-3/0/1.675 – MCS-240rtr – May experience a momentary blip of under 5 seconds
Xe-3/2/3 – CyberSpan - Will lose replicated traffic during the maintenance window (expected to be less than 10 minutes)
Xe-3/3/0.1300 – APS-ScienceDMZ – Will lose connectivity during the maintenance window (expected to be less than 10 minutes)
Please let me know if there are any issues or concerns. If there is a more suitable timeframe for the swap, please suggest.
corby
All:
At 0050hrs last night, the border router in 541b (noni) suffered a hardware-related event. Linecard #3 clocked a number of errors on traffic received from the chassis backplane. At 0057.16hrs, the chassis health monitor set the system status to Yellow, noting an impact to some traffic flows through linecard #3. At 0057.23hrs, the chassis health monitor set the system status to Red, indicating that the errors were unrecoverable in the current situation and that all traffic through linecard #3 was impacted. This culminated in a chassis health process initiating a module reset at 0057.38hrs, which ceased all traffic across linecard #3. At 0100.11hrs, linecard #3 had restored to operational state and by 0103.05hrs, all related network services had been fully restored.
The following links were impacted by this outage:
Xe-3/0/0.803,804 – MREN Backup Link (R&E and TR/CPS) – not active at the time of the event
Xe-3/0/0.345,346 – ESnet Tertiary Link – not active at the time of the event
Xe-3/0/1.675 – MCS-240rtr – networking on alternate path via lulo was not impacted, short outage of 15s was likely perceived
Xe-3/0/2.668 – T1-Firewall link deprecated during maintenance weekend – not in production state
Xe-3/1/1 – 308Core1-L2 – Layer 2 service backup link, all primary services were unaffected on 100G link to Core541
Xe-3/2/3 – CyberSpan, during this window the cyber monitoring on noni was non-functional – 0050-0103hrs
Xe-3/3/0.1300 – APS-ScienceDMZ – impacted throughout the window from 0050-0103hrs
We are working with Juniper TAC to identify the root cause of the hardware failure, and take whatever action needed to rectify the situation. There have been no additional indications of fabric errors since the card reset this morning, and no indication prior to the log messages at 0500hrs this morning.
More information will be provided as we move through the TAC process. If you have any questions or concerns, please reach out to me directly, or send a note to anchor@anl.gov.
--
Corby Schmitz
Network Communication Operations and Support Manager
Computer and Information Systems Division
-
Network Engineer
MREN/Starlight
--
Argonne National Laboratory
9700 S. Cass Ave.
Argonne, IL 60439
Desk: 630-252-7664
Cell: 630-296-4252
E-mail: cschmitz@anl.gov