I just want to check what the benchmark should do. My understanding from the conference calls and Sameer's slides is that they're interested in the case where multiple threads of a process are sending/receiving to/from different processes. This allows the process to use multiple DMAs concurrently. Right now, we've been testing the case where one process is sending/receiving with itself. So moving the locks around VCs may decrease the size of the CS, but will still serialize. So what's the test we want to run? I imagine something like P processes each with P-1 threads where each thread is communicating with a different process. But then we need a machine with P(P-1) cores. On the 8-core nodes we have, we can run 3 processes with 2 threads each. But I think we need more threads to really stress things. It seems that what we can see from the one or two processes is the just the effect of the size of the CS. Comments? -d