Ah, I had missed the fact that it didn't scale the results. So is the cost of the MPI call stack to the top of the ch3 code (where the proc-null test is) about 25ns on Breadboard and 60ns on BG/P?No. The latency number is already total_time / (LOOPS * 16), so you can't divide the latency number by 16 again.Sorry for the confusion.-- Pavan--Pavan Balaji