https://bugzilla.mcs.anl.gov/swift/show_bug.cgi?id=690 --- Comment #39 from Mihael Hategan <[email protected]> 2012-04-11 19:29:32 --- (In reply to comment #38)
Im adding thsi comment to document several emails describing progress in past 24 hours.
Mihael, is the fix to the scenario you describe in the first email below (Tuesday, April 10, 2012 7:04:56 PM) now comitted? David, is that what you tested? Or did this scenario not occur with the new NIO code in place?
Partially. There is still a deadlock possible which is mentioned in my email from Tuesday, April 10, 2012 7:04:56 PM. I just (as of a few hours ago) finished re-writing the worker socket read/write loop to use selectors and deal with the situation. I am currently testing this and ironing out the bugs. The NIO stuff is still something necessary to get maximum performance and prevent a single slow/blocked worker from slowing everything down.
Also, David: your email below implies that some of these fixes are not in place for automatic coasters, but only for manual (as you say "latest update which fixes coaster-service")
The fix should be there, but there was a bug present with manual coasters. David isn't seeing the code working because it really needs the updated worker which is not committed yet.
Can you both update this ticket to summarize the state of the fixes? Ie, are all the problems identified related to provider staging now fixed for all configs, or is there more known work needed to call this issue resolved?
As far as I can tell all the problems so far are identified and fixes are in the pipe. Though we won't know for sure until after these fixes are tested.
In any case, nice work guys - this is excellent news and excellent progress!
-- Configure bugmail: https://bugzilla.mcs.anl.gov/swift/userprefs.cgi?tab=email ------- You are receiving this mail because: ------- You are watching all bug changes.