BTW, the Breadboard results suggest that the timing program is too simple in this environment. It would be good for the timing program to run the same test multiple times and at least report the spread in the results. Do we even believe the 170ns latency time for the 1- thread case? If so, is that the new record? (Note also that the test violates the same-send/recv-buffer rule :) ) Bill On Jan 29, 2008, at 3:53 PM, Pavan Balaji wrote:
Hi all,
I've run some experiments for the MPICH2 threading overhead on BG/P and Breadboard (test program is attached). I don't remember if IBM reported that the performance stays constant with increasing threads or if it degrades, but I'm noticing a drop in performance (results below).
On BG/P, this is slight yet noticeable, but on breadboard it's drastic. Note that 1 thread refers to "no extra threads" -- the main process does a self send/recv.
This is not the fairest comparison since MPICH2-BG/P is based on MPICH2-1.0.4p1, while the breadboard runs are on MPICH2-trunk. I'll try out MPICH2-1.0.4p1 on breadboard as well. But as a longer term solution I'm trying to get MPICH2-BG/P ported to MPICH2-trunk so that I can easily try out any threading enhancements that might go in to trunk from here on.
Thanks.
-- Pavan
William Gropp Paul and Cynthia Saylor Professor of Computer Science University of Illinois Urbana-Champaign