I'm looking into the nagios alerts about the websites now. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
It looks like a network problem on the port channel between the mcs router and the mdf -- May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/4 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/5 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Po240 in err-disable state May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/4 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/5 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Po240 Interface Status Protocol Description Po240 up up t2-240mdf-PC:Core I left a message on corby's cell. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Core Admins" <[email protected]> Sent: Tuesday, May 24, 2011 11:03:58 PM Subject: web problems I'm looking into the nagios alerts about the websites now. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
Dan, Sorry for not taking the call, but I was in a deep sleep. It took an elbow from Sheri to bring me around and that was for the VM notification. All, There are many questions raised by this event: 1. why did this problem suddenly come back in the middle of the night whilst (I love that word) I was asleep and no engineers were active (this is a puzzler) If you remember, and likely you don't, this first happened when we added the VOIP network to the links up to the hackspace. It was fixed by adding a line to the configuration that told the switch to ignore spanning-tree, which is pretty dangerous in this king of L2 environment. After an upgrade of the MDF code, things got better and we turned spanning-tree back on toward the MDF. We have been running happily since then (roughly August of last year) 2. why did nagios not see the interface as down I suspect that this is a function of timing. Because the links were down it was down for exactly 1 minute at a time and the poling cycle is every 5 minutes, I guess we were just unlucky. Learning from this, I think its time to do instant alerts on events in the syslog stream coming out of the network gear. Its probably well overdue anyway. 3. why did this break servers at all Dan answered this one for me already ... its our fast-path to 221 where the load balancers live. The up and down flapping on regular intervals was driving the load balancers nuts and the servers paid the price. Based on the log messages, this seems to have started somewhere around 2220hrs this evening, but it may have taken a bit longer to begin to affect the servers. I am still digging through logs and trying to figure out the timeline. 4. why did I ever remove this line from the config: no spanning-tree etherchannel guard misconfig As I stated earlier, this is not a great thing to ignore as its a global command, but since the problem is back, I will be leaving it in place for the near term (at least till I have confirmed with both JTAC and CTAC that this is the only way to fix the interop problem). As a sidebar, the MDF switch upgrade on Friday was supposed to make this problem evaporate for CIS, so hopefully it does so for us as well. The good news is that things are stable at this point ... Nagios looks happy and servers are visible. I will follow up with JTAC and CTAC in the morning as CIS still has open cases on this issue related to the MDF switch to 221. I will post a note in the morning once I have a chance to sync with the TAC folks. -- Corby Schmitz Network Engineer CIS/MCS/MREN/Starlight Argonne National Laboratory 9700 S. Cass Ave. Argonne, IL 60439 Desk: 630-252-7664 Cell: 630-512-1502 E-mail: [email protected] On May 24, 2011, at 11:18 PM, Dan Olson wrote:
It looks like a network problem on the port channel between the mcs router and the mdf -- May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/4 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/5 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Po240 in err-disable state May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/4 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/5 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Po240
Interface Status Protocol Description Po240 up up t2-240mdf-PC:Core
I left a message on corby's cell.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Core Admins" <[email protected]> Sent: Tuesday, May 24, 2011 11:03:58 PM Subject: web problems
I'm looking into the nagios alerts about the websites now.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
Thanks, guys. -- Craig On May 25, 2011, at 12:00 AM, Schmitz Corby <[email protected]> wrote:
Dan, Sorry for not taking the call, but I was in a deep sleep. It took an elbow from Sheri to bring me around and that was for the VM notification.
All, There are many questions raised by this event: 1. why did this problem suddenly come back in the middle of the night whilst (I love that word) I was asleep and no engineers were active (this is a puzzler) If you remember, and likely you don't, this first happened when we added the VOIP network to the links up to the hackspace. It was fixed by adding a line to the configuration that told the switch to ignore spanning-tree, which is pretty dangerous in this king of L2 environment. After an upgrade of the MDF code, things got better and we turned spanning-tree back on toward the MDF. We have been running happily since then (roughly August of last year)
2. why did nagios not see the interface as down I suspect that this is a function of timing. Because the links were down it was down for exactly 1 minute at a time and the poling cycle is every 5 minutes, I guess we were just unlucky. Learning from this, I think its time to do instant alerts on events in the syslog stream coming out of the network gear. Its probably well overdue anyway.
3. why did this break servers at all Dan answered this one for me already ... its our fast-path to 221 where the load balancers live. The up and down flapping on regular intervals was driving the load balancers nuts and the servers paid the price. Based on the log messages, this seems to have started somewhere around 2220hrs this evening, but it may have taken a bit longer to begin to affect the servers. I am still digging through logs and trying to figure out the timeline.
4. why did I ever remove this line from the config: no spanning-tree etherchannel guard misconfig As I stated earlier, this is not a great thing to ignore as its a global command, but since the problem is back, I will be leaving it in place for the near term (at least till I have confirmed with both JTAC and CTAC that this is the only way to fix the interop problem). As a sidebar, the MDF switch upgrade on Friday was supposed to make this problem evaporate for CIS, so hopefully it does so for us as well.
The good news is that things are stable at this point ... Nagios looks happy and servers are visible. I will follow up with JTAC and CTAC in the morning as CIS still has open cases on this issue related to the MDF switch to 221. I will post a note in the morning once I have a chance to sync with the TAC folks.
-- Corby Schmitz Network Engineer CIS/MCS/MREN/Starlight Argonne National Laboratory 9700 S. Cass Ave. Argonne, IL 60439 Desk: 630-252-7664 Cell: 630-512-1502 E-mail: [email protected]
On May 24, 2011, at 11:18 PM, Dan Olson wrote:
It looks like a network problem on the port channel between the mcs router and the mdf -- May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/4 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/5 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Po240 in err-disable state May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/4 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/5 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Po240
Interface Status Protocol Description Po240 up up t2-240mdf-PC:Core
I left a message on corby's cell.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Core Admins" <[email protected]> Sent: Tuesday, May 24, 2011 11:03:58 PM Subject: web problems
I'm looking into the nagios alerts about the websites now.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
Hi, Guys, I'm not sure if this is related or not, but I experienced a lot of problems accessing the web server yesterday. At one point I even went in to Ken's office to see if it was down. Joe Insley also had problems connecting to his web space and at one point saw contents in his directory that should not have been there. Beth On May 25, 2011, at 12:25 AM, Craig Stacey wrote:
Thanks, guys.
-- Craig
On May 25, 2011, at 12:00 AM, Schmitz Corby <[email protected]> wrote:
Dan, Sorry for not taking the call, but I was in a deep sleep. It took an elbow from Sheri to bring me around and that was for the VM notification.
All, There are many questions raised by this event: 1. why did this problem suddenly come back in the middle of the night whilst (I love that word) I was asleep and no engineers were active (this is a puzzler) If you remember, and likely you don't, this first happened when we added the VOIP network to the links up to the hackspace. It was fixed by adding a line to the configuration that told the switch to ignore spanning-tree, which is pretty dangerous in this king of L2 environment. After an upgrade of the MDF code, things got better and we turned spanning-tree back on toward the MDF. We have been running happily since then (roughly August of last year)
2. why did nagios not see the interface as down I suspect that this is a function of timing. Because the links were down it was down for exactly 1 minute at a time and the poling cycle is every 5 minutes, I guess we were just unlucky. Learning from this, I think its time to do instant alerts on events in the syslog stream coming out of the network gear. Its probably well overdue anyway.
3. why did this break servers at all Dan answered this one for me already ... its our fast-path to 221 where the load balancers live. The up and down flapping on regular intervals was driving the load balancers nuts and the servers paid the price. Based on the log messages, this seems to have started somewhere around 2220hrs this evening, but it may have taken a bit longer to begin to affect the servers. I am still digging through logs and trying to figure out the timeline.
4. why did I ever remove this line from the config: no spanning- tree etherchannel guard misconfig As I stated earlier, this is not a great thing to ignore as its a global command, but since the problem is back, I will be leaving it in place for the near term (at least till I have confirmed with both JTAC and CTAC that this is the only way to fix the interop problem). As a sidebar, the MDF switch upgrade on Friday was supposed to make this problem evaporate for CIS, so hopefully it does so for us as well.
The good news is that things are stable at this point ... Nagios looks happy and servers are visible. I will follow up with JTAC and CTAC in the morning as CIS still has open cases on this issue related to the MDF switch to 221. I will post a note in the morning once I have a chance to sync with the TAC folks.
-- Corby Schmitz Network Engineer CIS/MCS/MREN/Starlight Argonne National Laboratory 9700 S. Cass Ave. Argonne, IL 60439 Desk: 630-252-7664 Cell: 630-512-1502 E-mail: [email protected]
On May 24, 2011, at 11:18 PM, Dan Olson wrote:
It looks like a network problem on the port channel between the mcs router and the mdf -- May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/4 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Te1/5 in err-disable state May 24 22:44:35 CDT: %PM-SP-4-ERR_DISABLE: channel-misconfig error detected on Po240, putting Po240 in err-disable state May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/4 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Te1/5 May 24 22:45:35 CDT: %PM-SP-4-ERR_RECOVER: Attempting to recover from channel-misconfig err-disable state on Po240
Interface Status Protocol Description Po240 up up t2-240mdf- PC:Core
I left a message on corby's cell.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Dan Olson" <[email protected]> To: "Core Admins" <[email protected]> Sent: Tuesday, May 24, 2011 11:03:58 PM Subject: web problems
I'm looking into the nagios alerts about the websites now.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
looking at the network issue on the MDF trunk now. -- Corby Schmitz Network Engineer CIS/MCS/MREN/Starlight Argonne National Laboratory 9700 S. Cass Ave. Argonne, IL 60439 Desk: 630-252-7664 Cell: 630-512-1502 E-mail: [email protected] On May 24, 2011, at 11:03 PM, Dan Olson wrote:
I'm looking into the nagios alerts about the websites now.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
participants (4)
-
Beth Cerny Patino -
Craig Stacey -
Dan Olson -
Schmitz Corby