Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it. On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Seems to be pretty unhappy still after the reboot. Looking more closely... Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites. -- Craig On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate. -- Craig On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Yes, we can easily throw one in the document root of both these sites. I'll do that today. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/04/2012 11:33 PM, Craig Stacey wrote:
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate.
-- Craig
On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Turns out Google ignores crawl rate directives in robots.txt. I'll have to configure this through the Google webmaster tools, which only last for 90 days. Should be a nice motivator to get variant on better hardware. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/05/2012 08:45 AM, Ken Raffenetti wrote:
Yes, we can easily throw one in the document root of both these sites. I'll do that today.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 11:33 PM, Craig Stacey wrote:
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate.
-- Craig
On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Okay, both trac.mcs and trac.alcf have been throttled to the lowest possible crawl rate through google's webmaster tools. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/05/2012 08:51 AM, Ken Raffenetti wrote:
Turns out Google ignores crawl rate directives in robots.txt. I'll have to configure this through the Google webmaster tools, which only last for 90 days. Should be a nice motivator to get variant on better hardware.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:45 AM, Ken Raffenetti wrote:
Yes, we can easily throw one in the document root of both these sites. I'll do that today.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 11:33 PM, Craig Stacey wrote:
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate.
-- Craig
On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/05/2012 08:59 AM, Ken Raffenetti wrote:
Okay, both trac.mcs and trac.alcf have been throttled to the lowest possible crawl rate through google's webmaster tools.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:51 AM, Ken Raffenetti wrote:
Turns out Google ignores crawl rate directives in robots.txt. I'll have to configure this through the Google webmaster tools, which only last for 90 days. Should be a nice motivator to get variant on better hardware.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:45 AM, Ken Raffenetti wrote:
Yes, we can easily throw one in the document root of both these sites. I'll do that today.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 11:33 PM, Craig Stacey wrote:
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate.
-- Craig
On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Can we move watcher to one of the barely used hypervisors? -- Craig On Jun 5, 2012, at 1:19 PM, Ken Raffenetti <[email protected]> wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:59 AM, Ken Raffenetti wrote:
Okay, both trac.mcs and trac.alcf have been throttled to the lowest possible crawl rate through google's webmaster tools.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:51 AM, Ken Raffenetti wrote:
Turns out Google ignores crawl rate directives in robots.txt. I'll have to configure this through the Google webmaster tools, which only last for 90 days. Should be a nice motivator to get variant on better hardware.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:45 AM, Ken Raffenetti wrote:
Yes, we can easily throw one in the document root of both these sites. I'll do that today.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 11:33 PM, Craig Stacey wrote:
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate.
-- Craig
On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Moving it is no problem. We'll need to take some downtime. It will take ~1 hour to move. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/05/2012 01:19 PM, Craig Stacey wrote:
Can we move watcher to one of the barely used hypervisors?
-- Craig
On Jun 5, 2012, at 1:19 PM, Ken Raffenetti <[email protected]> wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:59 AM, Ken Raffenetti wrote:
Okay, both trac.mcs and trac.alcf have been throttled to the lowest possible crawl rate through google's webmaster tools.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:51 AM, Ken Raffenetti wrote:
Turns out Google ignores crawl rate directives in robots.txt. I'll have to configure this through the Google webmaster tools, which only last for 90 days. Should be a nice motivator to get variant on better hardware.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:45 AM, Ken Raffenetti wrote:
Yes, we can easily throw one in the document root of both these sites. I'll do that today.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 11:33 PM, Craig Stacey wrote:
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate.
-- Craig
On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Let's do it this afternoon, no need to wait for off-hours. -- Craig (from my mobile) Ken Raffenetti <[email protected]> wrote: Moving it is no problem. We'll need to take some downtime. It will take ~1 hour to move. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/05/2012 01:19 PM, Craig Stacey wrote:
Can we move watcher to one of the barely used hypervisors?
-- Craig
On Jun 5, 2012, at 1:19 PM, Ken Raffenetti <[email protected]> wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:59 AM, Ken Raffenetti wrote:
Okay, both trac.mcs and trac.alcf have been throttled to the lowest possible crawl rate through google's webmaster tools.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:51 AM, Ken Raffenetti wrote:
Turns out Google ignores crawl rate directives in robots.txt. I'll have to configure this through the Google webmaster tools, which only last for 90 days. Should be a nice motivator to get variant on better hardware.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 08:45 AM, Ken Raffenetti wrote:
Yes, we can easily throw one in the document root of both these sites. I'll do that today.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 11:33 PM, Craig Stacey wrote:
Any way to put a robots.txt on trac.alcf.anl.gov and the other Trac sites? We can then set the crawl rate.
-- Craig
On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote:
Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites.
-- Craig
On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote:
Seems to be pretty unhappy still after the reboot. Looking more closely...
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/04/2012 09:40 AM, Ken Raffenetti wrote:
It was pretty unhappy this morning, so I kicked it. To add to the pain, I got a lock error when trying to restart the domain on vserver7.mcs.anl.gov. To remedy, I had to "/etc/init.d/libvirt-bin restart" before I could start the VM again.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/03/2012 06:01 PM, Craig Stacey wrote:
Whatever happened to it this evening, it seems to have fixed itself by the time I got to it.
On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote:
Dunno what was going on with variant -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. -- Craig
-- Craig
Shutting down watcher now. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/05/2012 01:25 PM, Craig Stacey wrote:
Let's do it this afternoon, no need to wait for off-hours. -- Craig (from my mobile)
Ken Raffenetti <[email protected]> wrote:
Moving it is no problem. We'll need to take some downtime. It will take ~1 hour to move.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:19 PM, Craig Stacey wrote: > Can we move watcher to one of the barely used hypervisors? > > -- > Craig > > On Jun 5, 2012, at 1:19 PM, Ken Raffenetti <[email protected]> wrote: > >> watcher is thrashing the disk on this hypervisor. That's what's causing >> variant to go into IO wait and spike load. >> >> Ken Raffenetti >> Systems Administration Associate >> MCS Division - Argonne National Laboratory >> >> On 06/05/2012 08:59 AM, Ken Raffenetti wrote: >>> Okay, both trac.mcs and trac.alcf have bee n throttled to the lowest >>> possible crawl rate through google's webmaster tools. >>> >>> Ken Raffenetti >>> Systems Administration Associate >>> MCS Division - Argonne National Laboratory >>> >>> On 06/05/2012 08:51 AM, Ken Raffenetti wrote: >>>> Turns out Google ignores crawl rate directives in robots.txt. I'll have >>>> to configure this through the Google webmaster tools, which only last >>>> for 90 days. Should be a nice motivator to get variant on better hardware. >>>> >>>> Ken Raffenetti >>>> Systems Administration Associate >>>> MCS Division - Argonne National Laboratory >>>> >>>> On 06/05/2012 08:45 AM, Ken Raffenetti wrote: >>>>> Yes, we can easily throw one in the document root of both these sites . >>>>> I'll do that today. >>>>> >>>>> Ken Raffenetti >>>>> Systems Administration Associate >>>>> MCS Division - Argonne National Laboratory >>>>> >>>>> On 06/04/2012 11:33 PM, Craig Stacey wrote: >>>>>> Any way to put a robots.txt on trac.alcf.anl.gov <http://trac.alcf.anl.gov> and the other Trac sites? We can then set the crawl rate. >>>>>> >>>>>> -- >>>>>> Craig >>>>>> >>>>>> On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote: >>>>>> >>>>>>> Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites. >>>>>>> >>>>>>& gt; -- >>>>>>> Craig >>>>>>> >>>>>>> On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote: >>>>>>> >>>>>>>> Seems to be pretty unhappy still after the reboot. Looking more closely... >>>>>>>> >>>>>>>> Ken Raffenetti >>>>>>>> Systems Administration Associate >>>>>>>> MCS Division - Argonne National Laboratory >>>>>>>> >>>>>>>> On 06/04/2012 09:40 AM, Ken Raffenetti wrote: >>>>>>>>> It was pretty unhappy this morning, so I kicked it. To add to the pain, >>>>>>>>> I got a lock error when trying to restart the domain on >>>>>>>>> vserver7.mcs.anl.gov <http://vserver7.mcs.anl.gov>. To remedy, I had to "/etc/init.d/libvirt-bin >>>>>>>>> restart" before I could start the VM again. >>>>>>>>> >>>>>>>>> Ken Raffenetti >>>>>>>>> Systems Administration Associate >>>>>>>>> MCS Division - Argonne National Laboratory >>>>>>>>> >>>>>>>>> On 06/03/2012 06:01 PM, Craig Stacey wrote: >>>>>>>>>> Whatever happened to it this evening, it seems to have fixed itself by the time I got to it. >>>>>>>>>> >>>>>>>>>> On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote: >>>>>>>>>> >>>>>>>>>>> Dunno what was going on with varian t -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. >>>>>>>>>>> -- >>>>>>>>>>> Craig >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>> >>>>>>>>>> -- >>>>>>>>>> Craig >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>> >>>>>>>>> >>>>>>>> >>>>>>>> >>>>> >>>>> >>>> >>>> >>> >>> >> >>
Noooooooooooooooooooooooooooooo! Just kidding. -- corby On 6/5/12 1:26 PM, "Raffenetti, Kenneth J." <[email protected]> wrote:
Shutting down watcher now.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:25 PM, Craig Stacey wrote:
Let's do it this afternoon, no need to wait for off-hours. -- Craig (from my mobile)
Ken Raffenetti <[email protected]> wrote:
Moving it is no problem. We'll need to take some downtime. It will take ~1 hour to move.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:19 PM, Craig Stacey wrote: > Can we move watcher to one of the barely used hypervisors? > > -- > Craig > > On Jun 5, 2012, at 1:19 PM, Ken Raffenetti <[email protected]> wrote: > >> watcher is thrashing the disk on this hypervisor. That's what's causing >> variant to go into IO wait and spike load. >> >> Ken Raffenetti >> Systems Administration Associate >> MCS Division - Argonne National Laboratory >> >> On 06/05/2012 08:59 AM, Ken Raffenetti wrote: >>> Okay, both trac.mcs and trac.alcf have bee n throttled to the lowest >>> possible crawl rate through google's webmaster tools. >>> >>> Ken Raffenetti >>> Systems Administration Associate >>> MCS Division - Argonne National Laboratory >>> >>> On 06/05/2012 08:51 AM, Ken Raffenetti wrote: >>>> Turns out Google ignores crawl rate directives in robots.txt. I'll have >>>> to configure this through the Google webmaster tools, which only last >>>> for 90 days. Should be a nice motivator to get variant on better hardware. >>>> >>>> Ken Raffenetti >>>> Systems Administration Associate >>>> MCS Division - Argonne National Laboratory >>>> >>>> On 06/05/2012 08:45 AM, Ken Raffenetti wrote: >>>>> Yes, we can easily throw one in the document root of both these sites . >>>>> I'll do that today. >>>>> >>>>> Ken Raffenetti >>>>> Systems Administration Associate >>>>> MCS Division - Argonne National Laboratory >>>>> >>>>> On 06/04/2012 11:33 PM, Craig Stacey wrote: >>>>>> Any way to put a robots.txt on trac.alcf.anl.gov <http://trac.alcf.anl.gov> and the other Trac sites? We can then set the crawl rate. >>>>>> >>>>>> -- >>>>>> Craig >>>>>> >>>>>> On Jun 4, 2012, at 11:29 PM, Craig Stacey <[email protected]> wrote: >>>>>> >>>>>>> Googlebot killed us again. Need to get that blocked until we figure out how to stop it from killing the Trac sites. >>>>>>> >>>>>>& gt; -- >>>>>>> Craig >>>>>>> >>>>>>> On Jun 4, 2012, at 9:45 AM, Ken Raffenetti <[email protected]> wrote: >>>>>>> >>>>>>>> Seems to be pretty unhappy still after the reboot. Looking more closely... >>>>>>>> >>>>>>>> Ken Raffenetti >>>>>>>> Systems Administration Associate >>>>>>>> MCS Division - Argonne National Laboratory >>>>>>>> >>>>>>>> On 06/04/2012 09:40 AM, Ken Raffenetti wrote: >>>>>>>>> It was pretty unhappy this morning, so I kicked it. To add to the pain, >>>>>>>>> I got a lock error when trying to restart the domain on >>>>>>>>> vserver7.mcs.anl.gov <http://vserver7.mcs.anl.gov>. To remedy, I had to "/etc/init.d/libvirt-bin >>>>>>>>> restart" before I could start the VM again. >>>>>>>>> >>>>>>>>> Ken Raffenetti >>>>>>>>> Systems Administration Associate >>>>>>>>> MCS Division - Argonne National Laboratory >>>>>>>>> >>>>>>>>> On 06/03/2012 06:01 PM, Craig Stacey wrote: >>>>>>>>>> Whatever happened to it this evening, it seems to have fixed itself by the time I got to it. >>>>>>>>>> >>>>>>>>>> On Jun 3, 2012, at 2:00 PM, Craig Stacey wrote: >>>>>>>>>> >>>>>>>>>>> Dunno what was going on with varian t -- load was through the roof. After various attempts to get services stopped and reduce load so I could do some actual troubleshooting, I decided it was taking too much time and just rebooted it. Things seemed to come back cleanly and it hasn't complained since. >>>>>>>>>>> -- >>>>>>>>>>> Craig >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>> >>>>>>>>>> -- >>>>>>>>>> Craig >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> >>>>>>>>> >>>>>>>>> >>>>>>>> >>>>>>>> >>>>> >>>>> >>>> >>>> >>> >>> >> >>
On Tue, Jun 05, 2012 at 01:19:06PM -0500, Ken Raffenetti wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Is watcher writing out of ton of Ganglia (or other) RRDs? If so, moving those to a tmpfs may make it happier. I had to do this w/ Ganglia on fusion, and the disk I/O dropped to nearly nothing as did the load. http://www.mail-archive.com/[email protected]/msg03060.h... John
This is a likely cause. Once we get watcher sectioned off we can experiment with this. Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory On 06/05/2012 01:49 PM, John Valdes wrote:
On Tue, Jun 05, 2012 at 01:19:06PM -0500, Ken Raffenetti wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Is watcher writing out of ton of Ganglia (or other) RRDs? If so, moving those to a tmpfs may make it happier. I had to do this w/ Ganglia on fusion, and the disk I/O dropped to nearly nothing as did the load. http://www.mail-archive.com/[email protected]/msg03060.h...
John
Right now, IIRC, the only things watcher is watching is temperature and power (which don't use Ganglia, I believe). Is it doing anything else that uses Ganglia? -- Craig On Jun 5, 2012, at 1:54 PM, Ken Raffenetti wrote:
This is a likely cause. Once we get watcher sectioned off we can experiment with this.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:49 PM, John Valdes wrote:
On Tue, Jun 05, 2012 at 01:19:06PM -0500, Ken Raffenetti wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Is watcher writing out of ton of Ganglia (or other) RRDs? If so, moving those to a tmpfs may make it happier. I had to do this w/ Ganglia on fusion, and the disk I/O dropped to nearly nothing as did the load. http://www.mail-archive.com/[email protected]/msg03060.h...
John
Watcher was doing ganglia monitoring in general on systems machines. We didn't stand up ganglia on peer. ---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055 ----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Ken Raffenetti" <[email protected]> Cc: "John Valdes" <[email protected]>, "core-a >> [email protected]" <[email protected]> Sent: Tuesday, June 5, 2012 1:55:58 PM Subject: Re: variant warnings Right now, IIRC, the only things watcher is watching is temperature and power (which don't use Ganglia, I believe). Is it doing anything else that uses Ganglia? -- Craig On Jun 5, 2012, at 1:54 PM, Ken Raffenetti wrote:
This is a likely cause. Once we get watcher sectioned off we can experiment with this.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:49 PM, John Valdes wrote:
On Tue, Jun 05, 2012 at 01:19:06PM -0500, Ken Raffenetti wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Is watcher writing out of ton of Ganglia (or other) RRDs? If so, moving those to a tmpfs may make it happier. I had to do this w/ Ganglia on fusion, and the disk I/O dropped to nearly nothing as did the load. http://www.mail-archive.com/[email protected]/msg03060.h...
John
I see, okay, that makes sense. -- Craig On Jun 5, 2012, at 2:03 PM, Dan Olson wrote:
Watcher was doing ganglia monitoring in general on systems machines. We didn't stand up ganglia on peer.
---- Daniel Murphy-Olson Systems Administrator Mathematics & Computer Science Division Argonne National Laboratory 630-252-0055
----- Original Message ----- From: "Craig Stacey" <[email protected]> To: "Ken Raffenetti" <[email protected]> Cc: "John Valdes" <[email protected]>, "core-a >> [email protected]" <[email protected]> Sent: Tuesday, June 5, 2012 1:55:58 PM Subject: Re: variant warnings
Right now, IIRC, the only things watcher is watching is temperature and power (which don't use Ganglia, I believe). Is it doing anything else that uses Ganglia? -- Craig
On Jun 5, 2012, at 1:54 PM, Ken Raffenetti wrote:
This is a likely cause. Once we get watcher sectioned off we can experiment with this.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:49 PM, John Valdes wrote:
On Tue, Jun 05, 2012 at 01:19:06PM -0500, Ken Raffenetti wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Is watcher writing out of ton of Ganglia (or other) RRDs? If so, moving those to a tmpfs may make it happier. I had to do this w/ Ganglia on fusion, and the disk I/O dropped to nearly nothing as did the load. http://www.mail-archive.com/[email protected]/msg03060.h...
John
No ganglia on watcher today. -- corby On 6/5/12 1:55 PM, "Stacey, Craig J." <[email protected]> wrote:
Right now, IIRC, the only things watcher is watching is temperature and power (which don't use Ganglia, I believe). Is it doing anything else that uses Ganglia? -- Craig
On Jun 5, 2012, at 1:54 PM, Ken Raffenetti wrote:
This is a likely cause. Once we get watcher sectioned off we can experiment with this.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:49 PM, John Valdes wrote:
On Tue, Jun 05, 2012 at 01:19:06PM -0500, Ken Raffenetti wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Is watcher writing out of ton of Ganglia (or other) RRDs? If so, moving those to a tmpfs may make it happier. I had to do this w/ Ganglia on fusion, and the disk I/O dropped to nearly nothing as did the load.
http://www.mail-archive.com/[email protected]/msg030 60.html
John
Note that other things use RRDs too (eg, mrtg), but ganglia in particular seems to be very efficient at generating RRDs. :) John On Tue, Jun 05, 2012 at 02:11:24PM -0500, Schmitz, Corby B. wrote:
No ganglia on watcher today.
-- corby
On 6/5/12 1:55 PM, "Stacey, Craig J." <[email protected]> wrote:
Right now, IIRC, the only things watcher is watching is temperature and power (which don't use Ganglia, I believe). Is it doing anything else that uses Ganglia? -- Craig
On Jun 5, 2012, at 1:54 PM, Ken Raffenetti wrote:
This is a likely cause. Once we get watcher sectioned off we can experiment with this.
Ken Raffenetti Systems Administration Associate MCS Division - Argonne National Laboratory
On 06/05/2012 01:49 PM, John Valdes wrote:
On Tue, Jun 05, 2012 at 01:19:06PM -0500, Ken Raffenetti wrote:
watcher is thrashing the disk on this hypervisor. That's what's causing variant to go into IO wait and spike load.
Is watcher writing out of ton of Ganglia (or other) RRDs? If so, moving those to a tmpfs may make it happier. I had to do this w/ Ganglia on fusion, and the disk I/O dropped to nearly nothing as did the load.
http://www.mail-archive.com/[email protected]/msg030 60.html
John
participants (5)
-
Craig Stacey -
Dan Olson -
John Valdes -
Ken Raffenetti -
Schmitz, Corby B.