29 May
2015
29 May
'15
2:56 p.m.
Mark Adams <[email protected]> writes:
Yea, I realized that VecAssembly should see this load imbalance unless it had a barrier before its timer. So I'm not sure what is going on.
Since you don't have any outgoing entries, I think the huge MPI_Allreduce is expensive. No surprise.