discuss
Threads by month
- ----- 2026 -----
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
August 2025
- 7 participants
- 16 discussions
Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
by Raffenetti, Ken 28 Aug '25
by Raffenetti, Ken 28 Aug '25
28 Aug '25
I also tried on Frontier and get the expected (good) latency results with the current main branch. The Slurm installation on Frontier does not support PMIx, so it is not quite apples to apples. Does your system have the Cray PMI library to try?
Ken
From: Raffenetti, Ken via discuss <discuss(a)mpich.org>
Date: Friday, August 22, 2025 at 11:03 AM
To: discuss(a)mpich.org <discuss(a)mpich.org>, Zhou, Hui <zhouh(a)anl.gov>
Cc: Raffenetti, Ken <raffenet(a)anl.gov>, discuss(a)mpich.org <discuss(a)mpich.org>
Subject: Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
Hi Howard,
The PMIx stuff is likely related to the new sessions implementation coming in 5.0.x. I’ll look for a Slurm cluster to try and figure out what’s going on with that.
What commit hash are you working with that shows the poor latency? I just built from the HEAD of main and don’t see the behavior on Aurora.
Ken
From: Howard Pritchard via discuss <discuss(a)mpich.org>
Date: Thursday, August 21, 2025 at 4:40 PM
To: Zhou, Hui <zhouh(a)anl.gov>
Cc: Howard Pritchard <hppritcha(a)gmail.com>, discuss(a)mpich.org <discuss(a)mpich.org>
Subject: Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
This Message Is From an External Sender
This message came from outside your organization.
Here you go Hui!
MPICH debug output and slurm steps output to boot. Again no such slurmy errors with the 4.3.1 release.
Something must have changed in the way MPICH is using the PMIX group constructor ops or something like that.
Required minimum FI_VERSION: 0, current version: 10016
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
Required minimum FI_VERSION: 10005, current version: 10016
==== Capability set configuration ====
libfabric provider: cxi - cxi
MPIDI_OFI_ENABLE_DATA: 1
MPIDI_OFI_ENABLE_AV_TABLE: 1
MPIDI_OFI_ENABLE_SCALABLE_ENDPOINTS: 0
MPIDI_OFI_ENABLE_SHARED_CONTEXTS: 0
MPIDI_OFI_ENABLE_MR_VIRT_ADDRESS: 0
MPIDI_OFI_ENABLE_MR_ALLOCATED: 1
MPIDI_OFI_ENABLE_MR_REGISTER_NULL: 0
MPIDI_OFI_ENABLE_MR_PROV_KEY: 0
MPIDI_OFI_ENABLE_TAGGED: 1
MPIDI_OFI_ENABLE_AM: 1
MPIDI_OFI_ENABLE_RMA: 1
MPIDI_OFI_ENABLE_ATOMICS: 1
MPIDI_OFI_FETCH_ATOMIC_IOVECS: 1
MPIDI_OFI_ENABLE_DATA_AUTO_PROGRESS: 0
MPIDI_OFI_ENABLE_CONTROL_AUTO_PROGRESS: 0
MPIDI_OFI_ENABLE_PT2PT_NOPACK: 1
MPIDI_OFI_ENABLE_TRIGGERED: 0
MPIDI_OFI_ENABLE_HMEM: 0
MPIDI_OFI_NUM_AM_BUFFERS: 8
MPIDI_OFI_NUM_OPTIMIZED_MEMORY_REGIONS: 0
MPIDI_OFI_CONTEXT_BITS: 20
MPIDI_OFI_SOURCE_BITS: 0
MPIDI_OFI_TAG_BITS: 20
MPIDI_OFI_VNI_USE_DOMAIN: 1
MAXIMUM SUPPORTED RANKS: 4294967296
MAXIMUM TAG: 1048576
==== Provider global thresholds ====
max_buffered_send: 192
max_buffered_write: 192
max_msg_size: 4294967295
max_order_raw: -1
max_order_war: -1
max_order_waw: -1
tx_iov_limit: 1
rx_iov_limit: 1
rma_iov_limit: 1
max_mr_key_size: 4
==== Various sizes and limits ====
MPIDI_OFI_AM_MSG_HEADER_SIZE: 24
MPIDI_OFI_MAX_AM_HDR_SIZE: 255
sizeof(MPIDI_OFI_am_request_header_t): 416
sizeof(MPIDI_OFI_per_vci_t): 52480
MPIDI_OFI_AM_HDR_POOL_CELL_SIZE: 1024
MPIDI_OFI_DEFAULT_SHORT_SEND_SIZE: 16384
======================================
slurmstepd: error: mpi/pmix_v4: pmixp_coll_belong_chk: nid001406 [1]: pmixp_coll.c:280: No process controlled by this slurmstepd is involved in this collective.
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001406 [1]: pmixp_server.c:923: Unable to pmixp_state_coll_get()
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_check: nid001405 [0]: pmixp_coll_ring.c:614: 0x14b448006e10: unexpected contrib from nid001406:1, expected is 0
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001405 [0]: pmixp_server.c:937: 0x14b448006e10: unexpected contrib from nid001406:1, coll->seq=0, seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001405 [0]: pmixp_coll_ring.c:738: 0x14b454052fc0: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001405 [0]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:756: 0x14b454052fc0: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:758: my peerid: 0:nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:765: neighbor id: next 1:nid001406, prev 1:nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:775: Context ptr=0x14b454053038, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:775: Context ptr=0x14b454053070, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:775: Context ptr=0x14b4540530a8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:823: wait contrib: nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001406 [1]: pmixp_coll_ring.c:738: 0x14aa28053100: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001406 [1]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:756: 0x14aa28053100: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:758: my peerid: 1:nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:765: neighbor id: next 0:nid001405, prev 0:nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:775: Context ptr=0x14aa28053178, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:775: Context ptr=0x14aa280531b0, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:775: Context ptr=0x14aa280531e8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:823: wait contrib: nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
==== Various sizes and limits ====
sizeof(MPIDI_per_vci_t): 128
==== collective selection ====
MPIR_CVAR_DEVICE_COLLECTIVES: percoll
MPIR: MPII_coll_generic_json
MPID: MPIDI_coll_generic_json
MPID/shm: MPIDI_POSIX_coll_generic_json
==== OFI dynamic settings ====
num_vcis: 1
num_nics: 1
======================================
error checking : disabled
QMPI : disabled
debugger support : disabled
thread level : MPI_THREAD_SINGLE
thread CS : per-vci
threadcomm : enabled
==== data structure summary ====
sizeof(MPIR_Comm): 1832
sizeof(MPIR_Request): 520
sizeof(MPIR_Datatype): 280
================================
# OSU MPI Latency Test v5.8
# Size Latency (us)
0 2.04
1 10.08
2 10.10
4 10.11
8 10.12
16 10.12
32 10.13
64 10.12
128 10.67
256 8.10
512 8.18
1024 8.11
2048 7.86
4096 7.80
8192 10.25
16384 11.04
32768 12.04
65536 14.05
131072 17.89
262144 24.61
524288 37.51
1048576 61.48
2097152 110.06
4194304 228.67
Am Mi., 13. Aug. 2025 um 13:10 Uhr schrieb Zhou, Hui <zhouh(a)anl.gov<mailto:[email protected]>>:
Hi Howard,
Could you run with `MPIR_CVAR_DEBUG_SUMMARY=1`? It should print some debug messages. I want to confirm it is running the `cxi` provider.
Hui
________________________________
From: Howard Pritchard <hppritcha(a)gmail.com<mailto:[email protected]>>
Sent: Wednesday, July 30, 2025 4:37 PM
To: Thakur, Rajeev <thakur(a)anl.gov<mailto:[email protected]>>
Cc: discuss(a)mpich.org<mailto:[email protected]> <discuss(a)mpich.org<mailto:[email protected]>>; Zhou, Hui <zhouh(a)anl.gov<mailto:[email protected]>>
Subject: Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
You don't often get email from hppritcha(a)gmail.com<mailto:[email protected]>. Learn why this is important<https://urldefense.us/v3/__https://aka.ms/LearnAboutSenderIdentification__;…>
This Message Is From an External Sender
This message came from outside your organization.
Hi Rajeev,
Here are the results for 4.3.x branch:
hpp@nid001293:/usr/projects/artab/users/hpp/osu-micro-benchmarks-5.8-mpich/mpi/pt2pt>srun --mpi=pmix -n 2 ./osu_latency
# OSU MPI Latency Test v5.8
# Size Latency (us)
0 1.92
1 1.98
2 1.98
4 1.98
8 1.98
16 1.98
32 1.99
64 1.99
128 2.47
256 2.59
512 2.65
1024 2.76
2048 2.95
4096 3.00
8192 5.96
16384 6.64
32768 7.44
65536 8.75
131072 11.52
262144 17.08
524288 27.96
1048576 49.38
2097152 92.96
4194304 179.74
These are more like i would expect for SS11/OFI CXI provider.
Howard
Am Mi., 30. Juli 2025 um 12:48 Uhr schrieb Thakur, Rajeev <thakur(a)anl.gov<mailto:[email protected]>>:
Hi Howard,
What was the latency with the 4.3.x branch?
Rajeev
From: Howard Pritchard via discuss <discuss(a)mpich.org<mailto:[email protected]>>
Reply-To: "discuss(a)mpich.org<mailto:[email protected]>" <discuss(a)mpich.org<mailto:[email protected]>>
Date: Wednesday, July 30, 2025 at 1:43 PM
To: "Zhou, Hui" <zhouh(a)anl.gov<mailto:[email protected]>>
Cc: Howard Pritchard <hppritcha(a)gmail.com<mailto:[email protected]>>, "discuss(a)mpich.org<mailto:[email protected]>" <discuss(a)mpich.org<mailto:[email protected]>>
Subject: Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
Hi Hui That didn’t help. I am not surprised though as our cluster is an NVIDIA free zone. What did help is to switch to the mpich 4. 3. x branch and latency results are nominal and the slurm problem went away too. So we will stick with that branch.
ZjQcmQRYFpfptBannerStart
This Message Is From an External Sender
This message came from outside your organization.
ZjQcmQRYFpfptBannerEnd
Hi Hui
That didn’t help. I am not surprised though as our cluster is an NVIDIA free zone. What did help is to switch to the mpich 4.3.x branch and latency results are nominal and the slurm problem went away too. So we will stick with that branch.
Howard
On Mon, Jul 28, 2025 at 4:15 PM Zhou, Hui <zhouh(a)anl.gov<mailto:[email protected]>> wrote:
Hi Howard,
I wonder whether it is due to the overhead of querying pointer attributes. Could you try disable GPU support with `MPIR_CVAR_ENABLE_GPU=0` and see if the latency improves?
Hui
________________________________
From: Howard Pritchard via discuss <discuss(a)mpich.org<mailto:[email protected]>>
Sent: Monday, July 28, 2025 9:41 AM
To: discuss(a)mpich.org<mailto:[email protected]> <discuss(a)mpich.org<mailto:[email protected]>>
Cc: Howard Pritchard <hppritcha(a)gmail.com<mailto:[email protected]>>
Subject: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
Hi Folks, We are seeing a strange performance issue on our HPE SS11 system when testing osu_latency inter-node with MPICH. First the info: system using libfabric 1. 22. 0 slurm - 24. 11. 5 Here's my mpichversion output: MPICH Version: 5. 0. 0a1
ZjQcmQRYFpfptBannerStart
This Message Is From an External Sender
This message came from outside your organization.
ZjQcmQRYFpfptBannerEnd
Hi Folks,
We are seeing a strange performance issue on our HPE SS11 system when testing osu_latency inter-node with MPICH.
First the info:
system using libfabric 1.22.0
slurm - 24.11.5
Here's my mpichversion output:
MPICH Version: 5.0.0a1
MPICH Release date: unreleased development copy
MPICH ABI: 0:0:0
MPICH Device: ch4:ofi
MPICH configure: --prefix=/XXXX/mpich_again/install --enable-g=no --enable-error-checking=no --with-device=ch4:ofi --enable-threads=multiple --with-ch4-shmmods=posix,xpmem --enable-thread-cs=per-vci --with-libfabric=/opt/cray/libfabric/1.22.0 --with-xpmem=/opt/cray/xpmem/default --with-pmix=/opt/pmix/gcc4x/5.0.8 --enable-fast=O3
MPICH CC: gcc -O3
MPICH CXX: g++ -O3
MPICH F77: gfortran -O3
MPICH FC: gfortran -O3
MPICH features: threadcomm
And here's the OSU latency results:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_belong_chk: nid001439 [1]: pmixp_coll.c:280: No process controlled by this slurmstepd is involved in this collective.
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001439 [1]: pmixp_server.c:923: Unable to pmixp_state_coll_get()
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_check: nid001438 [0]: pmixp_coll_ring.c:614: 0x15005c005dc0: unexpected contrib from nid001439:1, expected is 0
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001438 [0]: pmixp_server.c:937: 0x15005c005dc0: unexpected contrib from nid001439:1, coll->seq=0, seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001438 [0]: pmixp_coll_ring.c:738: 0x1500580532f0: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001438 [0]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:756: 0x1500580532f0: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:758: my peerid: 0:nid001438
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:765: neighbor id: next 1:nid001439, prev 1:nid001439
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:775: Context ptr=0x150058053368, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:775: Context ptr=0x1500580533a0, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:775: Context ptr=0x1500580533d8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:823: wait contrib: nid001439
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001439 [1]: pmixp_coll_ring.c:738: 0x151d0c053400: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001439 [1]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:756: 0x151d0c053400: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:758: my peerid: 1:nid001439
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:765: neighbor id: next 0:nid001438, prev 0:nid001438
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:775: Context ptr=0x151d0c053478, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:775: Context ptr=0x151d0c0534b0, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:775: Context ptr=0x151d0c0534e8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:823: wait contrib: nid001438
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
# OSU MPI Latency Test v5.8
# Size Latency (us)
0 1.66
1 9.29
2 9.57
4 9.69
8 9.76
16 9.77
32 9.76
64 9.77
128 10.32
256 7.54
512 7.45
1024 7.38
2048 7.37
4096 7.45
8192 9.21
16384 9.70
32768 10.63
65536 13.15
131072 16.96
262144 23.84
524288 36.16
1048576 60.36
2097152 108.43
4194304 228.31
Note the slurm behavior is - I launch the job. Go get coffee, do some duo-lingo, read some emails, then after about 10 minutes the osu latency runs.
I did not get the slurm problems using an older mpich 4.3.1 but did get the same performance issue. 9 usecs doesn't seem right for an 8-byte pingpong over libfabric S11. I was expecting more like 1.6 or so.
I am confident the slurm issue is unrelated to the latency issue.
Thanks for any suggestions on how to address either issue however.
1
0
Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
by Raffenetti, Ken 22 Aug '25
by Raffenetti, Ken 22 Aug '25
22 Aug '25
Hi Howard,
The PMIx stuff is likely related to the new sessions implementation coming in 5.0.x. I’ll look for a Slurm cluster to try and figure out what’s going on with that.
What commit hash are you working with that shows the poor latency? I just built from the HEAD of main and don’t see the behavior on Aurora.
Ken
From: Howard Pritchard via discuss <discuss(a)mpich.org>
Date: Thursday, August 21, 2025 at 4:40 PM
To: Zhou, Hui <zhouh(a)anl.gov>
Cc: Howard Pritchard <hppritcha(a)gmail.com>, discuss(a)mpich.org <discuss(a)mpich.org>
Subject: Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
This Message Is From an External Sender
This message came from outside your organization.
Here you go Hui!
MPICH debug output and slurm steps output to boot. Again no such slurmy errors with the 4.3.1 release.
Something must have changed in the way MPICH is using the PMIX group constructor ops or something like that.
Required minimum FI_VERSION: 0, current version: 10016
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
Required minimum FI_VERSION: 10005, current version: 10016
==== Capability set configuration ====
libfabric provider: cxi - cxi
MPIDI_OFI_ENABLE_DATA: 1
MPIDI_OFI_ENABLE_AV_TABLE: 1
MPIDI_OFI_ENABLE_SCALABLE_ENDPOINTS: 0
MPIDI_OFI_ENABLE_SHARED_CONTEXTS: 0
MPIDI_OFI_ENABLE_MR_VIRT_ADDRESS: 0
MPIDI_OFI_ENABLE_MR_ALLOCATED: 1
MPIDI_OFI_ENABLE_MR_REGISTER_NULL: 0
MPIDI_OFI_ENABLE_MR_PROV_KEY: 0
MPIDI_OFI_ENABLE_TAGGED: 1
MPIDI_OFI_ENABLE_AM: 1
MPIDI_OFI_ENABLE_RMA: 1
MPIDI_OFI_ENABLE_ATOMICS: 1
MPIDI_OFI_FETCH_ATOMIC_IOVECS: 1
MPIDI_OFI_ENABLE_DATA_AUTO_PROGRESS: 0
MPIDI_OFI_ENABLE_CONTROL_AUTO_PROGRESS: 0
MPIDI_OFI_ENABLE_PT2PT_NOPACK: 1
MPIDI_OFI_ENABLE_TRIGGERED: 0
MPIDI_OFI_ENABLE_HMEM: 0
MPIDI_OFI_NUM_AM_BUFFERS: 8
MPIDI_OFI_NUM_OPTIMIZED_MEMORY_REGIONS: 0
MPIDI_OFI_CONTEXT_BITS: 20
MPIDI_OFI_SOURCE_BITS: 0
MPIDI_OFI_TAG_BITS: 20
MPIDI_OFI_VNI_USE_DOMAIN: 1
MAXIMUM SUPPORTED RANKS: 4294967296
MAXIMUM TAG: 1048576
==== Provider global thresholds ====
max_buffered_send: 192
max_buffered_write: 192
max_msg_size: 4294967295
max_order_raw: -1
max_order_war: -1
max_order_waw: -1
tx_iov_limit: 1
rx_iov_limit: 1
rma_iov_limit: 1
max_mr_key_size: 4
==== Various sizes and limits ====
MPIDI_OFI_AM_MSG_HEADER_SIZE: 24
MPIDI_OFI_MAX_AM_HDR_SIZE: 255
sizeof(MPIDI_OFI_am_request_header_t): 416
sizeof(MPIDI_OFI_per_vci_t): 52480
MPIDI_OFI_AM_HDR_POOL_CELL_SIZE: 1024
MPIDI_OFI_DEFAULT_SHORT_SEND_SIZE: 16384
======================================
slurmstepd: error: mpi/pmix_v4: pmixp_coll_belong_chk: nid001406 [1]: pmixp_coll.c:280: No process controlled by this slurmstepd is involved in this collective.
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001406 [1]: pmixp_server.c:923: Unable to pmixp_state_coll_get()
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_check: nid001405 [0]: pmixp_coll_ring.c:614: 0x14b448006e10: unexpected contrib from nid001406:1, expected is 0
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001405 [0]: pmixp_server.c:937: 0x14b448006e10: unexpected contrib from nid001406:1, coll->seq=0, seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001405 [0]: pmixp_coll_ring.c:738: 0x14b454052fc0: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001405 [0]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:756: 0x14b454052fc0: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:758: my peerid: 0:nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:765: neighbor id: next 1:nid001406, prev 1:nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:775: Context ptr=0x14b454053038, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:775: Context ptr=0x14b454053070, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:775: Context ptr=0x14b4540530a8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:823: wait contrib: nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001406 [1]: pmixp_coll_ring.c:738: 0x14aa28053100: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001406 [1]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:756: 0x14aa28053100: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:758: my peerid: 1:nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:765: neighbor id: next 0:nid001405, prev 0:nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:775: Context ptr=0x14aa28053178, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:775: Context ptr=0x14aa280531b0, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:775: Context ptr=0x14aa280531e8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:823: wait contrib: nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
==== Various sizes and limits ====
sizeof(MPIDI_per_vci_t): 128
==== collective selection ====
MPIR_CVAR_DEVICE_COLLECTIVES: percoll
MPIR: MPII_coll_generic_json
MPID: MPIDI_coll_generic_json
MPID/shm: MPIDI_POSIX_coll_generic_json
==== OFI dynamic settings ====
num_vcis: 1
num_nics: 1
======================================
error checking : disabled
QMPI : disabled
debugger support : disabled
thread level : MPI_THREAD_SINGLE
thread CS : per-vci
threadcomm : enabled
==== data structure summary ====
sizeof(MPIR_Comm): 1832
sizeof(MPIR_Request): 520
sizeof(MPIR_Datatype): 280
================================
# OSU MPI Latency Test v5.8
# Size Latency (us)
0 2.04
1 10.08
2 10.10
4 10.11
8 10.12
16 10.12
32 10.13
64 10.12
128 10.67
256 8.10
512 8.18
1024 8.11
2048 7.86
4096 7.80
8192 10.25
16384 11.04
32768 12.04
65536 14.05
131072 17.89
262144 24.61
524288 37.51
1048576 61.48
2097152 110.06
4194304 228.67
Am Mi., 13. Aug. 2025 um 13:10 Uhr schrieb Zhou, Hui <zhouh(a)anl.gov<mailto:[email protected]>>:
Hi Howard,
Could you run with `MPIR_CVAR_DEBUG_SUMMARY=1`? It should print some debug messages. I want to confirm it is running the `cxi` provider.
Hui
________________________________
From: Howard Pritchard <hppritcha(a)gmail.com<mailto:[email protected]>>
Sent: Wednesday, July 30, 2025 4:37 PM
To: Thakur, Rajeev <thakur(a)anl.gov<mailto:[email protected]>>
Cc: discuss(a)mpich.org<mailto:[email protected]> <discuss(a)mpich.org<mailto:[email protected]>>; Zhou, Hui <zhouh(a)anl.gov<mailto:[email protected]>>
Subject: Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
You don't often get email from hppritcha(a)gmail.com<mailto:[email protected]>. Learn why this is important<https://urldefense.us/v3/__https://aka.ms/LearnAboutSenderIdentification__;…>
This Message Is From an External Sender
This message came from outside your organization.
Hi Rajeev,
Here are the results for 4.3.x branch:
hpp@nid001293:/usr/projects/artab/users/hpp/osu-micro-benchmarks-5.8-mpich/mpi/pt2pt>srun --mpi=pmix -n 2 ./osu_latency
# OSU MPI Latency Test v5.8
# Size Latency (us)
0 1.92
1 1.98
2 1.98
4 1.98
8 1.98
16 1.98
32 1.99
64 1.99
128 2.47
256 2.59
512 2.65
1024 2.76
2048 2.95
4096 3.00
8192 5.96
16384 6.64
32768 7.44
65536 8.75
131072 11.52
262144 17.08
524288 27.96
1048576 49.38
2097152 92.96
4194304 179.74
These are more like i would expect for SS11/OFI CXI provider.
Howard
Am Mi., 30. Juli 2025 um 12:48 Uhr schrieb Thakur, Rajeev <thakur(a)anl.gov<mailto:[email protected]>>:
Hi Howard,
What was the latency with the 4.3.x branch?
Rajeev
From: Howard Pritchard via discuss <discuss(a)mpich.org<mailto:[email protected]>>
Reply-To: "discuss(a)mpich.org<mailto:[email protected]>" <discuss(a)mpich.org<mailto:[email protected]>>
Date: Wednesday, July 30, 2025 at 1:43 PM
To: "Zhou, Hui" <zhouh(a)anl.gov<mailto:[email protected]>>
Cc: Howard Pritchard <hppritcha(a)gmail.com<mailto:[email protected]>>, "discuss(a)mpich.org<mailto:[email protected]>" <discuss(a)mpich.org<mailto:[email protected]>>
Subject: Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
Hi Hui That didn’t help. I am not surprised though as our cluster is an NVIDIA free zone. What did help is to switch to the mpich 4. 3. x branch and latency results are nominal and the slurm problem went away too. So we will stick with that branch.
ZjQcmQRYFpfptBannerStart
This Message Is From an External Sender
This message came from outside your organization.
ZjQcmQRYFpfptBannerEnd
Hi Hui
That didn’t help. I am not surprised though as our cluster is an NVIDIA free zone. What did help is to switch to the mpich 4.3.x branch and latency results are nominal and the slurm problem went away too. So we will stick with that branch.
Howard
On Mon, Jul 28, 2025 at 4:15 PM Zhou, Hui <zhouh(a)anl.gov<mailto:[email protected]>> wrote:
Hi Howard,
I wonder whether it is due to the overhead of querying pointer attributes. Could you try disable GPU support with `MPIR_CVAR_ENABLE_GPU=0` and see if the latency improves?
Hui
________________________________
From: Howard Pritchard via discuss <discuss(a)mpich.org<mailto:[email protected]>>
Sent: Monday, July 28, 2025 9:41 AM
To: discuss(a)mpich.org<mailto:[email protected]> <discuss(a)mpich.org<mailto:[email protected]>>
Cc: Howard Pritchard <hppritcha(a)gmail.com<mailto:[email protected]>>
Subject: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
Hi Folks, We are seeing a strange performance issue on our HPE SS11 system when testing osu_latency inter-node with MPICH. First the info: system using libfabric 1. 22. 0 slurm - 24. 11. 5 Here's my mpichversion output: MPICH Version: 5. 0. 0a1
ZjQcmQRYFpfptBannerStart
This Message Is From an External Sender
This message came from outside your organization.
ZjQcmQRYFpfptBannerEnd
Hi Folks,
We are seeing a strange performance issue on our HPE SS11 system when testing osu_latency inter-node with MPICH.
First the info:
system using libfabric 1.22.0
slurm - 24.11.5
Here's my mpichversion output:
MPICH Version: 5.0.0a1
MPICH Release date: unreleased development copy
MPICH ABI: 0:0:0
MPICH Device: ch4:ofi
MPICH configure: --prefix=/XXXX/mpich_again/install --enable-g=no --enable-error-checking=no --with-device=ch4:ofi --enable-threads=multiple --with-ch4-shmmods=posix,xpmem --enable-thread-cs=per-vci --with-libfabric=/opt/cray/libfabric/1.22.0 --with-xpmem=/opt/cray/xpmem/default --with-pmix=/opt/pmix/gcc4x/5.0.8 --enable-fast=O3
MPICH CC: gcc -O3
MPICH CXX: g++ -O3
MPICH F77: gfortran -O3
MPICH FC: gfortran -O3
MPICH features: threadcomm
And here's the OSU latency results:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_belong_chk: nid001439 [1]: pmixp_coll.c:280: No process controlled by this slurmstepd is involved in this collective.
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001439 [1]: pmixp_server.c:923: Unable to pmixp_state_coll_get()
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_check: nid001438 [0]: pmixp_coll_ring.c:614: 0x15005c005dc0: unexpected contrib from nid001439:1, expected is 0
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001438 [0]: pmixp_server.c:937: 0x15005c005dc0: unexpected contrib from nid001439:1, coll->seq=0, seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001438 [0]: pmixp_coll_ring.c:738: 0x1500580532f0: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001438 [0]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:756: 0x1500580532f0: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:758: my peerid: 0:nid001438
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:765: neighbor id: next 1:nid001439, prev 1:nid001439
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:775: Context ptr=0x150058053368, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:775: Context ptr=0x1500580533a0, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:775: Context ptr=0x1500580533d8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:823: wait contrib: nid001439
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001439 [1]: pmixp_coll_ring.c:738: 0x151d0c053400: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001439 [1]: pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:756: 0x151d0c053400: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:758: my peerid: 1:nid001439
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:765: neighbor id: next 0:nid001438, prev 0:nid001438
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:775: Context ptr=0x151d0c053478, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:775: Context ptr=0x151d0c0534b0, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:775: Context ptr=0x151d0c0534e8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:823: wait contrib: nid001438
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]: pmixp_coll_ring.c:829: buf (offset/size): 36/16384
# OSU MPI Latency Test v5.8
# Size Latency (us)
0 1.66
1 9.29
2 9.57
4 9.69
8 9.76
16 9.77
32 9.76
64 9.77
128 10.32
256 7.54
512 7.45
1024 7.38
2048 7.37
4096 7.45
8192 9.21
16384 9.70
32768 10.63
65536 13.15
131072 16.96
262144 23.84
524288 36.16
1048576 60.36
2097152 108.43
4194304 228.31
Note the slurm behavior is - I launch the job. Go get coffee, do some duo-lingo, read some emails, then after about 10 minutes the osu latency runs.
I did not get the slurm problems using an older mpich 4.3.1 but did get the same performance issue. 9 usecs doesn't seem right for an 8-byte pingpong over libfabric S11. I was expecting more like 1.6 or so.
I am confident the slurm issue is unrelated to the latency issue.
Thanks for any suggestions on how to address either issue however.
1
0
Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more - a slurm problem
by Howard Pritchard 21 Aug '25
by Howard Pritchard 21 Aug '25
21 Aug '25
Here you go Hui!
MPICH debug output and slurm steps output to boot. Again no such
slurmy errors with the 4.3.1 release.
Something must have changed in the way MPICH is using the PMIX group
constructor ops or something like that.
Required minimum FI_VERSION: 0, current version: 10016
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
provider: cxi, score = 5, pref = 0, FI_FORMAT_UNSPEC [8]
Required minimum FI_VERSION: 10005, current version: 10016
==== Capability set configuration ====
libfabric provider: cxi - cxi
MPIDI_OFI_ENABLE_DATA: 1
MPIDI_OFI_ENABLE_AV_TABLE: 1
MPIDI_OFI_ENABLE_SCALABLE_ENDPOINTS: 0
MPIDI_OFI_ENABLE_SHARED_CONTEXTS: 0
MPIDI_OFI_ENABLE_MR_VIRT_ADDRESS: 0
MPIDI_OFI_ENABLE_MR_ALLOCATED: 1
MPIDI_OFI_ENABLE_MR_REGISTER_NULL: 0
MPIDI_OFI_ENABLE_MR_PROV_KEY: 0
MPIDI_OFI_ENABLE_TAGGED: 1
MPIDI_OFI_ENABLE_AM: 1
MPIDI_OFI_ENABLE_RMA: 1
MPIDI_OFI_ENABLE_ATOMICS: 1
MPIDI_OFI_FETCH_ATOMIC_IOVECS: 1
MPIDI_OFI_ENABLE_DATA_AUTO_PROGRESS: 0
MPIDI_OFI_ENABLE_CONTROL_AUTO_PROGRESS: 0
MPIDI_OFI_ENABLE_PT2PT_NOPACK: 1
MPIDI_OFI_ENABLE_TRIGGERED: 0
MPIDI_OFI_ENABLE_HMEM: 0
MPIDI_OFI_NUM_AM_BUFFERS: 8
MPIDI_OFI_NUM_OPTIMIZED_MEMORY_REGIONS: 0
MPIDI_OFI_CONTEXT_BITS: 20
MPIDI_OFI_SOURCE_BITS: 0
MPIDI_OFI_TAG_BITS: 20
MPIDI_OFI_VNI_USE_DOMAIN: 1
MAXIMUM SUPPORTED RANKS: 4294967296
MAXIMUM TAG: 1048576
==== Provider global thresholds ====
max_buffered_send: 192
max_buffered_write: 192
max_msg_size: 4294967295
max_order_raw: -1
max_order_war: -1
max_order_waw: -1
tx_iov_limit: 1
rx_iov_limit: 1
rma_iov_limit: 1
max_mr_key_size: 4
==== Various sizes and limits ====
MPIDI_OFI_AM_MSG_HEADER_SIZE: 24
MPIDI_OFI_MAX_AM_HDR_SIZE: 255
sizeof(MPIDI_OFI_am_request_header_t): 416
sizeof(MPIDI_OFI_per_vci_t): 52480
MPIDI_OFI_AM_HDR_POOL_CELL_SIZE: 1024
MPIDI_OFI_DEFAULT_SHORT_SEND_SIZE: 16384
======================================
slurmstepd: error: mpi/pmix_v4: pmixp_coll_belong_chk: nid001406 [1]:
pmixp_coll.c:280: No process controlled by this slurmstepd is involved in
this collective.
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001406 [1]:
pmixp_server.c:923: Unable to pmixp_state_coll_get()
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_check: nid001405 [0]:
pmixp_coll_ring.c:614: 0x14b448006e10: unexpected contrib from nid001406:1,
expected is 0
slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001405 [0]:
pmixp_server.c:937: 0x14b448006e10: unexpected contrib from nid001406:1,
coll->seq=0, seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001405
[0]: pmixp_coll_ring.c:738: 0x14b454052fc0: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001405 [0]:
pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:756: 0x14b454052fc0: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:758: my peerid: 0:nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:765: neighbor id: next 1:nid001406, prev 1:nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:775: Context ptr=0x14b454053038, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:775: Context ptr=0x14b454053070, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:775: Context ptr=0x14b4540530a8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:823: wait contrib: nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001405 [0]:
pmixp_coll_ring.c:829: buf (offset/size): 36/16384
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001406
[1]: pmixp_coll_ring.c:738: 0x14aa28053100: collective timeout seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001406 [1]:
pmixp_coll.c:286: Dumping collective state
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:756: 0x14aa28053100: COLL_FENCE_RING state seq=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:758: my peerid: 1:nid001406
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:765: neighbor id: next 0:nid001405, prev 0:nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:775: Context ptr=0x14aa28053178, #0, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:775: Context ptr=0x14aa280531b0, #1, in-use=0
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:775: Context ptr=0x14aa280531e8, #2, in-use=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:788: neighbor contribs [2]:
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:821: done contrib: -
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:823: wait contrib: nid001405
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001406 [1]:
pmixp_coll_ring.c:829: buf (offset/size): 36/16384
==== Various sizes and limits ====
sizeof(MPIDI_per_vci_t): 128
==== collective selection ====
MPIR_CVAR_DEVICE_COLLECTIVES: percoll
MPIR: MPII_coll_generic_json
MPID: MPIDI_coll_generic_json
MPID/shm: MPIDI_POSIX_coll_generic_json
==== OFI dynamic settings ====
num_vcis: 1
num_nics: 1
======================================
error checking : disabled
QMPI : disabled
debugger support : disabled
thread level : MPI_THREAD_SINGLE
thread CS : per-vci
threadcomm : enabled
==== data structure summary ====
sizeof(MPIR_Comm): 1832
sizeof(MPIR_Request): 520
sizeof(MPIR_Datatype): 280
================================
# OSU MPI Latency Test v5.8
# Size Latency (us)
0 2.04
1 10.08
2 10.10
4 10.11
8 10.12
16 10.12
32 10.13
64 10.12
128 10.67
256 8.10
512 8.18
1024 8.11
2048 7.86
4096 7.80
8192 10.25
16384 11.04
32768 12.04
65536 14.05
131072 17.89
262144 24.61
524288 37.51
1048576 61.48
2097152 110.06
4194304 228.67
Am Mi., 13. Aug. 2025 um 13:10 Uhr schrieb Zhou, Hui <zhouh(a)anl.gov>:
> Hi Howard,
>
> Could you run with `MPIR_CVAR_DEBUG_SUMMARY=1`? It should print some debug
> messages. I want to confirm it is running the `cxi` provider.
>
>
> Hui
> ------------------------------
> *From:* Howard Pritchard <hppritcha(a)gmail.com>
> *Sent:* Wednesday, July 30, 2025 4:37 PM
> *To:* Thakur, Rajeev <thakur(a)anl.gov>
> *Cc:* discuss(a)mpich.org <discuss(a)mpich.org>; Zhou, Hui <zhouh(a)anl.gov>
> *Subject:* Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus
> more - a slurm problem
>
> You don't often get email from hppritcha(a)gmail.com. Learn why this is
> important <https://urldefense.us/v3/__https://aka.ms/LearnAboutSenderIdentification__;… >
> Hi Rajeev, Here are the results for 4. 3. x branch: hpp@ nid001293:
> /usr/projects/artab/users/hpp/osu-micro-benchmarks-5.
> 8-mpich/mpi/pt2pt>srun --mpi=pmix -n 2 ./osu_latency # OSU MPI Latency Test
> v5. 8 # Size Latency (us) 0 1. 92 1 1. 98 2 1. 98
> ZjQcmQRYFpfptBannerStart
> This Message Is From an External Sender
> This message came from outside your organization.
>
> ZjQcmQRYFpfptBannerEnd
> Hi Rajeev,
>
> Here are the results for 4.3.x branch:
>
> hpp@nid001293:/usr/projects/artab/users/hpp/osu-micro-benchmarks-5.8-mpich/mpi/pt2pt>srun
> --mpi=pmix -n 2 ./osu_latency
>
> # OSU MPI Latency Test v5.8
>
> # Size Latency (us)
>
> 0 1.92
>
> 1 1.98
>
> 2 1.98
>
> 4 1.98
>
> 8 1.98
>
> 16 1.98
>
> 32 1.99
>
> 64 1.99
>
> 128 2.47
>
> 256 2.59
>
> 512 2.65
>
> 1024 2.76
>
> 2048 2.95
>
> 4096 3.00
>
> 8192 5.96
>
> 16384 6.64
>
> 32768 7.44
>
> 65536 8.75
>
> 131072 11.52
>
> 262144 17.08
>
> 524288 27.96
>
> 1048576 49.38
>
> 2097152 92.96
>
> 4194304 179.74
>
> These are more like i would expect for SS11/OFI CXI provider.
>
> Howard
>
> Am Mi., 30. Juli 2025 um 12:48 Uhr schrieb Thakur, Rajeev <thakur(a)anl.gov
> >:
>
> Hi Howard,
>
> What was the latency with the 4.3.x branch?
>
>
>
> Rajeev
>
>
>
>
>
> *From: *Howard Pritchard via discuss <discuss(a)mpich.org>
> *Reply-To: *"discuss(a)mpich.org" <discuss(a)mpich.org>
> *Date: *Wednesday, July 30, 2025 at 1:43 PM
> *To: *"Zhou, Hui" <zhouh(a)anl.gov>
> *Cc: *Howard Pritchard <hppritcha(a)gmail.com>, "discuss(a)mpich.org" <
> discuss(a)mpich.org>
> *Subject: *Re: [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus
> more - a slurm problem
>
>
>
> Hi Hui That didn’t help. I am not surprised though as our cluster is an
> NVIDIA free zone. What did help is to switch to the mpich 4. 3. x branch
> and latency results are nominal and the slurm problem went away too. So we
> will stick with that branch.
>
> ZjQcmQRYFpfptBannerStart
>
> *This Message Is From an External Sender *
>
> This message came from outside your organization.
>
> ZjQcmQRYFpfptBannerEnd
>
> Hi Hui
>
>
>
> That didn’t help. I am not surprised though as our cluster is an NVIDIA
> free zone. What did help is to switch to the mpich 4.3.x branch and
> latency results are nominal and the slurm problem went away too. So we
> will stick with that branch.
>
>
>
> Howard
>
>
>
> On Mon, Jul 28, 2025 at 4:15 PM Zhou, Hui <zhouh(a)anl.gov> wrote:
>
> Hi Howard,
>
>
>
> I wonder whether it is due to the overhead of querying pointer
> attributes. Could you try disable GPU support with `MPIR_CVAR_ENABLE_GPU=0`
> and see if the latency improves?
>
>
>
> Hui
> ------------------------------
>
> *From:* Howard Pritchard via discuss <discuss(a)mpich.org>
> *Sent:* Monday, July 28, 2025 9:41 AM
> *To:* discuss(a)mpich.org <discuss(a)mpich.org>
> *Cc:* Howard Pritchard <hppritcha(a)gmail.com>
> *Subject:* [mpich-discuss] MPICH 5.0.1 performance on HPE SS11 plus more
> - a slurm problem
>
>
>
> Hi Folks, We are seeing a strange performance issue on our HPE SS11 system
> when testing osu_latency inter-node with MPICH. First the info: system
> using libfabric 1. 22. 0 slurm - 24. 11. 5 Here's my mpichversion output:
> MPICH Version: 5. 0. 0a1
>
> ZjQcmQRYFpfptBannerStart
>
> *This Message Is From an External Sender *
>
> This message came from outside your organization.
>
>
>
> ZjQcmQRYFpfptBannerEnd
>
> Hi Folks,
>
>
>
> We are seeing a strange performance issue on our HPE SS11 system when
> testing osu_latency inter-node with MPICH.
>
>
>
> First the info:
>
> system using libfabric 1.22.0
>
> slurm - 24.11.5
>
>
>
> Here's my mpichversion output:
>
>
>
> MPICH Version: 5.0.0a1
>
> MPICH Release date: unreleased development copy
>
> MPICH ABI: 0:0:0
>
> MPICH Device: ch4:ofi
>
> MPICH configure: --prefix=/XXXX/mpich_again/install --enable-g=no
> --enable-error-checking=no --with-device=ch4:ofi --enable-threads=multiple
> --with-ch4-shmmods=posix,xpmem --enable-thread-cs=per-vci
> --with-libfabric=/opt/cray/libfabric/1.22.0
> --with-xpmem=/opt/cray/xpmem/default --with-pmix=/opt/pmix/gcc4x/5.0.8
> --enable-fast=O3
>
> MPICH CC: gcc -O3
>
> MPICH CXX: g++ -O3
>
> MPICH F77: gfortran -O3
>
> MPICH FC: gfortran -O3
>
> MPICH features: threadcomm
>
>
>
> And here's the OSU latency results:
>
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_belong_chk: nid001439 [1]:
> pmixp_coll.c:280: No process controlled by this slurmstepd is involved in
> this collective.
>
> slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001439 [1]:
> pmixp_server.c:923: Unable to pmixp_state_coll_get()
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_check: nid001438 [0]:
> pmixp_coll_ring.c:614: 0x15005c005dc0: unexpected contrib from nid001439:1,
> expected is 0
>
> slurmstepd: error: mpi/pmix_v4: _process_server_request: nid001438 [0]:
> pmixp_server.c:937: 0x15005c005dc0: unexpected contrib from nid001439:1,
> coll->seq=0, seq=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001438
> [0]: pmixp_coll_ring.c:738: 0x1500580532f0: collective timeout seq=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001438 [0]:
> pmixp_coll.c:286: Dumping collective state
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:756: 0x1500580532f0: COLL_FENCE_RING state seq=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:758: my peerid: 0:nid001438
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:765: neighbor id: next 1:nid001439, prev 1:nid001439
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:775: Context ptr=0x150058053368, #0, in-use=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:775: Context ptr=0x1500580533a0, #1, in-use=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:775: Context ptr=0x1500580533d8, #2, in-use=1
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:788: neighbor contribs [2]:
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:821: done contrib: -
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:823: wait contrib: nid001439
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001438 [0]:
> pmixp_coll_ring.c:829: buf (offset/size): 36/16384
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_reset_if_to: nid001439
> [1]: pmixp_coll_ring.c:738: 0x151d0c053400: collective timeout seq=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_log: nid001439 [1]:
> pmixp_coll.c:286: Dumping collective state
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:756: 0x151d0c053400: COLL_FENCE_RING state seq=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:758: my peerid: 1:nid001439
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:765: neighbor id: next 0:nid001438, prev 0:nid001438
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:775: Context ptr=0x151d0c053478, #0, in-use=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:775: Context ptr=0x151d0c0534b0, #1, in-use=0
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:775: Context ptr=0x151d0c0534e8, #2, in-use=1
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:786: seq=0 contribs: loc=1/prev=0/fwd=1
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:788: neighbor contribs [2]:
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:821: done contrib: -
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:823: wait contrib: nid001438
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:825: status=PMIXP_COLL_RING_PROGRESS
>
> slurmstepd: error: mpi/pmix_v4: pmixp_coll_ring_log: nid001439 [1]:
> pmixp_coll_ring.c:829: buf (offset/size): 36/16384
>
> # OSU MPI Latency Test v5.8
>
> # Size Latency (us)
>
> 0 1.66
>
> 1 9.29
>
> 2 9.57
>
> 4 9.69
>
> 8 9.76
>
> 16 9.77
>
> 32 9.76
>
> 64 9.77
>
> 128 10.32
>
> 256 7.54
>
> 512 7.45
>
> 1024 7.38
>
> 2048 7.37
>
> 4096 7.45
>
> 8192 9.21
>
> 16384 9.70
>
> 32768 10.63
>
> 65536 13.15
>
> 131072 16.96
>
> 262144 23.84
>
> 524288 36.16
>
> 1048576 60.36
>
> 2097152 108.43
>
> 4194304 228.31
>
>
>
> Note the slurm behavior is - I launch the job. Go get coffee, do some
> duo-lingo, read some emails, then after about 10 minutes the osu latency
> runs.
>
>
>
> I did not get the slurm problems using an older mpich 4.3.1 but did get
> the same performance issue. 9 usecs doesn't seem right for an 8-byte
> pingpong over libfabric S11. I was expecting more like 1.6 or so.
>
>
>
> I am confident the slurm issue is unrelated to the latency issue.
>
> Thanks for any suggestions on how to address either issue however.
>
>
>
>
>
>
1
0
False alarm. I did a cleanup and started over from scratch. Now,
everything's fine. Thanks anyway.
Regards.
1
0
1
0
1
0
Hi Sam,
> The main takeaways are that the single- and multi-core memory bandwidth appears to be the same for both machines, although the IPC bandwidth observed using MPI is >1.5x higher on the single-socket machine than the original machine that is showing the slow-down.
That seems to say that intra-socket bandwidth is higher than inter-socket bandwidth. That is typical. However, 1.5x seems extreme. Because the single core bandwidth bottleneck should be the CPU pipeline bandwidth rather than the package memory bandwidth. Nevertheless, 3.5GB bandwidth does seem low. It should reach ~10GB/s. Something is slowing it down and I am not sure I have a clue yet.
Just to double check, could you try the following (not at the same time) -
1. Set MPIR_CVAR_ODD_EVEN_CLIQUES=1
2. Set MPIR_CVAR_CH4_CMA_ENABLE=1
Also I think you made a typo in the last email. Please confirm that you are using MPICH-4.3.1, not MPICH-4.2.1, right?
Hui
________________________________
From: Sam Austin via discuss <discuss(a)mpich.org>
Sent: Thursday, August 14, 2025 7:16 PM
To: Boyle, Peter <pboyle(a)bnl.gov>
Cc: Sam Austin <sam.austin.p(a)gmail.com>; discuss(a)mpich.org <discuss(a)mpich.org>
Subject: Re: [mpich-discuss] MPICH: SHM bandwidth very low on IPC test
Hi all, Thank you for your advice; I really appreciate it. Try repeat the measurement a few times, you should see higher bandwidth number in the later rounds. I did this by running the transfer 100 times and only using the last 20 measurements
ZjQcmQRYFpfptBannerStart
This Message Is From an External Sender
This message came from outside your organization.
ZjQcmQRYFpfptBannerEnd
Hi all,
Thank you for your advice; I really appreciate it.
Try repeat the measurement a few times, you should see higher bandwidth number in the later rounds.
I did this by running the transfer 100 times and only using the last 20 measurements (discarding the first 80), and I saw the same performance.
Obviously if there are multiple rank pairs on the node communicating concurrently, then multiple cores will be active copying concurrently increasing aggregate bandwidth to closer to the many core or threaded throughput, which is what you would get if you compile STREAM with "-fopenmp”.
Good point -- I ran stream.c with -fopenmp, which returned 48 GB/s.
So, the discrepancy seems to be when dealing with IPC. Next, I ran these benchmarks, also using MPICH 4.2.1, on a different machine from the same era (Xeon E5-2650 v4). This is a single-socket board with one processor, as opposed to the original Dell Poweredge in question, which has two sockets, each with a E5-2699A. For both machines, I have attached the following in a zip file to avoid clogging the thread:
* Hardware topology (from lstopo)
* lscpu output
* Output from stream.c (single core)
* Output from stream_omp.c (openmp)
* Output from my host-host communication program (shmem_check.cpp):
The main takeaways are that the single- and multi-core memory bandwidth appears to be the same for both machines, although the IPC bandwidth observed using MPI is >1.5x higher on the single-socket machine than the original machine that is showing the slow-down.
I would assume, that your MPI bandwidth calculation only accounts for the buffer size (i.e., only read or write).
This is a good point, although I'm not sure why the bandwidth would be 1.5x higher on the other machine that has very similar memory performance.
One more piece of information: The single-socket machine (Xeon E5-2650 v4) has four RAM sticks in a quad-channel configuration, all tied to the same socket as you can see in lstopo. On the dual-socket machine (the machine in question), the four RAM sticks are in a dual-channel configuration, with two sticks on each socket. So, I'm not sure if the dual- vs quad-channel configuration is hurting maximum memory bandwidth per socket on the dual-socket machine, despite the total bandwidths showing approximately the same in STREAM.
I have tried to make the tests more uniform by binding the processes to the same core. But that still calls into question the total memory bandwidth of the dual-channel vs quad-channel memory configuration.
Do you have any thoughts on this? The question I'm trying to answer is: for what reasons would the memory bandwidth be significantly lower on the dual-socket Dell machine? I understand that making comparisons across machines is tricky, but I've tried to provide as much information as possible to isolate the key aspects of the memory configuration.
Thanks in advance!
Sam
On Thu, Aug 14, 2025 at 12:41 PM Boyle, Peter <pboyle(a)bnl.gov<mailto:[email protected]>> wrote:
Hi,
1)
Obviously if there are multiple rank pairs on the node communicating
concurrently, then multiple cores will be active copying concurrently increasing
aggregate bandwidth to closer to the many core or threaded throughput, which is what you
would get if you compile STREAM with "-fopenmp”.
The Xeon part you quote has a peak of 76GB/s memory
bandwidth per socket and multiple cores will be needed to saturate that.
2) It is common for people to use hybrid OpenMP and MPI, to minimize intra-node
copy overhead. Often using one rank per NUMA domain.
In that context, using MPI_Comm_split(…, MPI_COMM_TYPE_SHARED) will reveal which ranks can
use unix shared memory regions (e.g. shmopen ) and then use ALL their threads concurrently
to blast data between sockets, and use OpenMP within a socket.
Then we get into the topic of careful NUMA binding of ranks to sockets etc...
Best wishes,
Peter
From: Joachim Jenke via discuss <discuss(a)mpich.org<mailto:[email protected]>>
Date: Thursday, August 14, 2025 at 12:21 PM
To: Sam Austin <sam.austin.p(a)gmail.com<mailto:[email protected]>>
Cc: Joachim Jenke <jenke(a)itc.rwth-aachen.de<mailto:[email protected]>>, discuss(a)mpich.org<mailto:[email protected]> <discuss(a)mpich.org<mailto:[email protected]>>
Subject: Re: [mpich-discuss] MPICH: SHM bandwidth very low on IPC test
Hi Sam,
the 10GB/s stream bandwidth calculation includes the number of
read+written bytes (see lines 190/366).
I would assume, that your MPI bandwidth calculation only accounts for
the buffer size (i.e., only read or write). In shm communication one
process (and therefore one core) streams/memcopies the data from the
send to the receive buffer. So, when you see 3.5GB send bandwidth, that
actually compares to 7GB of stream Copy bandwidth.
As a side-effect of shm communication, we have actually seen that the
placement of the copying process can determine the first-touch
allocation of the buffer. Even if the memory is allocated with calloc,
the memory is not paged. A bcast/scatter to node-local processes can
result in paging all buffers to the same socket (what you typically want
to avoid).
Best
Joachim
Am 14.08.25 um 07:16 schrieb Sam Austin:
> Hi Joachim,
>
> Thanks for this suggestion! I used stream to test the single-core memory
> bandwidth. I am running on a Xeon E5-2699A v4, which has 55MB last level
> cache. So, I ran with 30 million elements per the instructions. It
> appears that I am seeing about 10 GB/s if I'm reading that right? If so,
> I am still not sure why I am only seeing ~3.5 GB/s on shared memory
> performance with MPICH.
>
> -------------------------------------------------------------
> STREAM version $Revision: 5.10 $
> -------------------------------------------------------------
> This system uses 8 bytes per array element.
> -------------------------------------------------------------
> Array size = 30000000 (elements), Offset = 0 (elements)
> Memory per array = 228.9 MiB (= 0.2 GiB).
> Total memory required = 686.6 MiB (= 0.7 GiB).
> Each kernel will be executed 10 times.
> The *best* time for each kernel (excluding the first iteration)
> will be used to compute the reported bandwidth.
> -------------------------------------------------------------
> Your clock granularity/precision appears to be 1 microseconds.
> Each test below will take on the order of 30276 microseconds.
> (= 30276 clock ticks)
> Increase the size of the arrays if this shows that
> you are not getting at least 20 clock ticks per test.
> -------------------------------------------------------------
> WARNING -- The above is only a rough guideline.
> For best results, please be sure you know the
> precision of your system timer.
> -------------------------------------------------------------
> Function Best Rate MB/s Avg time Min time Max time
> Copy: 10038.1 0.048342 0.047818 0.050004
> Scale: 10342.2 0.048738 0.046412 0.056605
> Add: 10580.3 0.068542 0.068051 0.069805
> Triad: 10703.0 0.067615 0.067271 0.068143
> -------------------------------------------------------------
> Solution Validates: avg error less than 1.000000e-13 on all three arrays
> -------------------------------------------------------------
>
> Thanks,
> Sam
>
> On Wed, Aug 13, 2025 at 5:26 PM Jenke, Joachim <jenke(a)itc.rwth-aachen.de<mailto:[email protected]>
> <mailto:[email protected]>> wrote:
>
> Hi Sam,
>
> Can you try out stream to understand the single-core memory
> bandwidth of the system?
>
> https://urldefense.us/v3/__https://www.cs.virginia.edu/stream/ref.html__;!!… <https://urldefense.us/v3/__https://www.cs.virginia.edu/stream/ref.html__;!!…> <https://
> https://urldefense.us/v3/__http://www.cs.virginia.edu/stream/ref.html__;!!G… ><https://urldefense.us/v3/__http://www.cs.virginia.edu/stream/ref.html*3E__;…>
>
> Copy bandwidth for large junks (exceeding cache sizes) should
> provide you an upper bound for shm communication bandwidth.
>
> Best
> Joachim
>
> Am 13.08.2025 22:04 schrieb Sam Austin via discuss
> <discuss(a)mpich.org<mailto:[email protected]> <mailto:[email protected]>>:
> Hi all, I am working to configure MPICH and run a few examples on my
> standalone server (single node). Here are the system specs: Server:
> Dell PowerEdge C4130 CPUs: 2x Xeon E5-2699A v4 GPUs: 4x Tesla V100s
> connected with NVLink, tied to motherboard
> ZjQcmQRYFpfptBannerStart
> This Message Is From an External Sender
> This message came from outside your organization.
> ZjQcmQRYFpfptBannerEnd
> Hi all,
>
> I am working to configure MPICH and run a few examples on my
> standalone server (single node). Here are the system specs:
> Server: Dell PowerEdge C4130
> CPUs: 2x Xeon E5-2699A v4
> GPUs: 4x Tesla V100s connected with NVLink, tied to motherboard with
> PCIe gen 3
> OS: Ubuntu 24.04 LTS
> I intend to use this system to develop multi-process programs for
> eventual execution in a large, distributed HPC environment. I ran a
> few tests with and without CUDA support; here is my mpichversion output:
>
> MPICH Version: 4.3.1
> MPICH Release date: Fri Jun 20 09:24:41 AM CDT 2025
> MPICH ABI: 17:1:5
> MPICH Device: ch4:ofi
> MPICH configure: --prefix=/opt/mpich/4.2.1-cpu --without-cuda
> MPICH CC: gcc -O2
> MPICH CXX: g++ -O2
> MPICH F77: gfortran -O2
> MPICH FC: gfortran -O2
> MPICH features: threadcomm
>
> The first example that I ran was a bandwidth test for CPU-CPU and
> GPU-GPU communication. This simple program sends small packets back
> and forth between processes to test the bandwidth over the various
> intra-node networks.
>
> The GPU-GPU bandwidth test showed that the GPU interconnect was
> saturating at ~45 GB/s, which is nominal for the NVLink interconnect
> topology present on the node (this was run with a CUDA-aware build
> of MPICH). The problem appears during the CPU-CPU IPC test. In
> theory, this test is pretty vanilla, as it is communicating between
> processes using shared memory, and does not involve traversing any
> of the intra-node networks (PCIe or NVLink). My understanding is
> that the bandwidth observed on the CPU-CPU IPC test should be quite
> high, at least higher than 10 GB/s.
>
> However, the intra-node IPC bandwidth appears to be very low, around
> 3.5 GB/s, when running this test. I tried the following fixes in an
> attempt to force MPICH to use shared memory, but to no avail:
> Passing the option to explicitly specify `nemesis` during the build
> configuration: "--with-device=ch3:nemesis --with-cuda"
> Passing the option to explicitly specify shared memory with ch4 to
> the configuration: "--with-ch4-shmmods=posix --with-cuda"
> Rebuilding MPICH without GPU support: "--without-cuda"
> Switching to Open MPI and running the same test
> These results, especially the last one in which I saw the same
> issues when running with Open MPI, makes me think it might be an
> issue with my system configuration. The question is: why is the IPC
> bandwidth so low despite supposedly using the SHM protocol? I'm
> wondering if anyone has encountered this issue before or might be
> able to lend some advice here. Any help would be greatly appreciated!
>
> Some interesting observations from the output below: when I run with
> "mpiexec -np 2 -genv FI_PROVIDER=shm ...", the log file reports
> "Opened fabric: shm". However, when I run without "-genv
> FI_PROVIDER=shm", the log file reports "Opened fabric: 10.133.0.0/21<https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!aD-lRdCn7h…>
> <https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!
> ZaUD7Nw-
> pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> kgoP1Dp-C6$>", which I believe means that MPICH is falling back on
> the TCP socket protocol. In this case, my key point of confusion is
> that the observed bandwidth is essentially the same between the SHM
> and TCP protocols. Perhaps my test script isn't set up properly?
>
> Thanks,
> Sam
>
> The following is attached below:
> Bandwidth test program
> Run script for the program
> Output of the script on my machine
> ----------------------------------------------------------------------------------------------------------------
> In case the attachment doesn't go through, here are the contents of
> my test program, "shmem_check.cpp":
>
> // shmem_check.cpp
> //
> // This is a minimal benchmark to test the raw bandwidth of MPI
> communication
> // between two processes on the same node, using only host (CPU) memory.
> // It completely removes CUDA to isolate the performance of the MPI
> library's
> // on-node communication mechanism (e.g., shared memory vs. TCP
> loopback).
> //
> // Compile/run:
> // /opt/mpich/4.2.1-cpu/bin/mpicxx -std=c++17 -I/opt/mpich/4.2.1-
> cpu/include shmem_check.cpp -o shmem_check
> // /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 ./shmem_check
>
> #include <iostream>
> #include <vector>
> #include <numeric>
> #include <mpi.h>
>
> int main(int argc, char* argv[]) {
> MPI_Init(&argc, &argv);
>
> int rank, size;
> MPI_Comm_rank(MPI_COMM_WORLD, &rank);
> MPI_Comm_size(MPI_COMM_WORLD, &size);
>
> if (size != 2) {
> if (rank == 0) {
> std::cerr << "Error: This program must be run with
> exactly 2 MPI processes." << std::endl;
> }
> MPI_Finalize();
> return 1;
> }
>
> const int num_samples = 100;
> const long long packet_size = 1LL << 28; // 256 MB
>
> // Allocate standard host memory. 'new' is sufficient.
> char* buffer = new char[packet_size];
>
> if (rank == 0) {
> std::cout << "--- Starting Host-to-Host MPI Bandwidth Test
> ---" << std::endl;
> std::cout << "Packet Size: " << (packet_size / (1024*1024))
> << " MB" << std::endl;
> }
>
> std::vector<double> timings;
> for (int i = 0; i < num_samples; ++i) {
> MPI_Barrier(MPI_COMM_WORLD);
> double start_time = MPI_Wtime();
>
> if (rank == 0) {
> MPI_Send(buffer, packet_size, MPI_CHAR, 1, 0,
> MPI_COMM_WORLD);
> MPI_Recv(buffer, 1, MPI_CHAR, 1, 1, MPI_COMM_WORLD,
> MPI_STATUS_IGNORE); // Wait for confirmation
> } else { // rank == 1
> MPI_Recv(buffer, packet_size, MPI_CHAR, 0, 0,
> MPI_COMM_WORLD, MPI_STATUS_IGNORE);
> MPI_Send(buffer, 1, MPI_CHAR, 0, 1, MPI_COMM_WORLD); //
> Send confirmation
> }
>
> double end_time = MPI_Wtime();
> if (i >= 10) { // Discard warmup runs
> timings.push_back(end_time - start_time);
> }
> }
>
> if (rank == 0) {
> double total_time = std::accumulate(timings.begin(),
> timings.end(), 0.0);
> double avg_time = total_time / timings.size();
> double bandwidth = (static_cast<double>(packet_size) /
> (1024.0 * 1024.0 * 1024.0)) / avg_time;
>
> std::cout <<
> "------------------------------------------------" << std::endl;
> std::cout << "Average Host-to-Host Bandwidth: " <<
> bandwidth << " GB/s" << std::endl;
> std::cout <<
> "------------------------------------------------" << std::endl;
> }
>
> // Clean up host memory
> delete[] buffer;
>
> MPI_Finalize();
> return 0;
> }
>
> ----------------------------------------------------------------------------------------------------------------
> Here is the script to run the test with verbose compilation and the
> `shm` layer forced and unforced:
>
> #!/usr/bin/zsh
> source ~/.zshrc
>
> # Compile
> /opt/mpich/4.2.1-cpu/bin/mpicxx -std=c++17 -I/opt/mpich/4.2.1-cpu/
> include shmem_check.cpp -o shmem_check
>
> # Run with shm forced
> /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 -genv FI_PROVIDER=shm -genv
> FI_LOG_LEVEL=debug ./shmem_check 2> output_shm.txt
>
> # Run without shm forced
> /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 -genv FI_LOG_LEVEL=debug ./
> shmem_check 2> output_no_shm.txt
>
> echo "Output of script with SHM forced: "
> grep -i "opened fabric" output_shm.txt
>
> echo "Output of script with SHM not forced: "
> grep -i "opened fabric" output_no_shm.txt
>
> ----------------------------------------------------------------------------------------------------------------
> Here is the output :
>
> --- Starting Host-to-Host MPI Bandwidth Test ---
> Packet Size: 256 MB
> ------------------------------------------------
> Average Host-to-Host Bandwidth: 3.35709 GB/s
> ------------------------------------------------
> --- Starting Host-to-Host MPI Bandwidth Test ---
> Packet Size: 256 MB
> ------------------------------------------------
> Average Host-to-Host Bandwidth: 3.54924 GB/s
> ------------------------------------------------
> Output of script with SHM forced:
> libfabric:3174297:1755114546::core:core:fi_fabric_():1503<info>
> Opened fabric: shm
> libfabric:3174298:1755114546::core:core:fi_fabric_():1503<info>
> Opened fabric: shm
> Output of script with SHM not forced:
> libfabric:3174351:1755114554::core:core:fi_fabric_():1503<info>
> Opened fabric: 10.133.0.0/21<https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!aD-lRdCn7h…> <https://urldefense.us/v3/
> __https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__;!!G_uCfscf7eWS!c1XQlmy02FA2ml32-W10DRZBMXRNNQHQbWR_bFh2y7B_3RgKnGGj_-cD3dOBBkOTUGUz6KdFvJY9$ <https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__…>
> pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> kgoP1Dp-C6$>
> libfabric:3174350:1755114554::core:core:fi_fabric_():1503<info>
> Opened fabric: 10.133.0.0/21<https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!aD-lRdCn7h…> <https://urldefense.us/v3/
> __https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__;!!G_uCfscf7eWS!c1XQlmy02FA2ml32-W10DRZBMXRNNQHQbWR_bFh2y7B_3RgKnGGj_-cD3dOBBkOTUGUz6KdFvJY9$ <https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__…>
> pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> kgoP1Dp-C6$>
>
--
Dr. rer. nat. Joachim Jenke
Deputy Group Lead
IT Center
Group: HPC - Parallelism, Runtime Analysis & Machine Learning
Division: Computational Science and Engineering
RWTH Aachen University
Seffenter Weg 23
D 52074 Aachen (Germany)
Tel: +49 241 80- 24765
Fax: +49 241 80-624765
jenke(a)itc.rwth-aachen.de<mailto:[email protected]>
https://urldefense.us/v3/__http://www.itc.rwth-aachen.de__;!!G_uCfscf7eWS!c… <https://urldefense.us/v3/__http://www.itc.rwth-aachen.de__;!!G_uCfscf7eWS!a…>
1
0
15 Aug '25
Hi Sam,
First I'm not sure why reaching top performance on a dev system would matter. To understand the production performance, you would need to test on the target system anyways. Especially because a single-node system is no good proxy for communication in a distributed memory system.
Am 15.08.2025 02:16 schrieb Sam Austin <sam.austin.p(a)gmail.com>:
I would assume, that your MPI bandwidth calculation only accounts for the buffer size (i.e., only read or write).
This is a good point, although I'm not sure why the bandwidth would be 1.5x higher on the other machine that has very similar memory performance.
Depending on the process placement, this might be caused by moving the data from one to the other socket, see below.
One more piece of information: The single-socket machine (Xeon E5-2650 v4) has four RAM sticks in a quad-channel configuration, all tied to the same socket as you can see in lstopo. On the dual-socket machine (the machine in question), the four RAM sticks are in a dual-channel configuration, with two sticks on each socket. So, I'm not sure if the dual- vs quad-channel configuration is hurting maximum memory bandwidth per socket on the dual-socket machine, despite the total bandwidths showing approximately the same in STREAM.
As Peter already pointed out: there are at least two bandwidth limits in a multicore system. One is the bandwidth a core can stream using it's load/store pipelines. The other bandwidth is for the connection between the CPU package and the memory. Going from Dual-Channel to Quad-Channel you increase the latter bandwidth. Therefore you will need more processes/threads to saturate the bandwidth with Quad-Channel configuration.
In Multi-Socket systems you have a third bandwidth limit for accessing memory connected with a different socket, which is typically lower. For the p2p bandwidth test, the bandwidth should not be the limiting factor, but the increased latency of these accesses might cause a reduced single-core bandwidth.
I have tried to make the tests more uniform by binding the processes to the same core. But that still calls into question the total memory bandwidth of the dual-channel vs quad-channel memory configuration.
Do you mean both processes to the same core, or symmetric proc placement on the two systems?
Try binding the processes to cores on the same/different sockets. Also, make sure to initialize the buffers before starting communication, so that they are paged locally. Repeated communication in the same direction might cause the OS to trigger page migration. So make sure to communicate back and forth.
Do you have any thoughts on this? The question I'm trying to answer is: for what reasons would the memory bandwidth be significantly lower on the dual-socket Dell machine? I understand that making comparisons across machines is tricky, but I've tried to provide as much information as possible to isolate the key aspects of the memory configuration.
At the end, p2p bandwidth will never be reached in a large-scale program, because filling the node will quickly saturate the overall memory bandwidth.
Best
Joachim
1
0
Hi all,
Thank you for your advice; I really appreciate it.
*Try repeat the measurement a few times, you should see higher bandwidth
number in the later rounds.*
I did this by running the transfer 100 times and only using the last 20
measurements (discarding the first 80), and I saw the same performance.
*Obviously if there are multiple rank pairs on the node communicating
concurrently, then multiple cores will be active copying concurrently
increasing aggregate bandwidth to closer to the many core or threaded
throughput, which is what you would get if you compile STREAM with
"-fopenmp”. *
Good point -- I ran stream.c with -fopenmp, which returned 48 GB/s.
So, the discrepancy seems to be when dealing with IPC. Next, I ran these
benchmarks, also using MPICH 4.2.1, on a different machine from the same
era (Xeon E5-2650 v4). This is a single-socket board with one processor, as
opposed to the original Dell Poweredge in question, which has two sockets,
each with a E5-2699A. For both machines, I have attached the following in a
zip file to avoid clogging the thread:
- Hardware topology (from lstopo)
- lscpu output
- Output from stream.c (single core)
- Output from stream_omp.c (openmp)
- Output from my host-host communication program (shmem_check.cpp):
The main takeaways are that the single- and multi-core memory bandwidth
appears to be the same for both machines, although the IPC bandwidth
observed using MPI is >*1.5x higher on the single-socket machine* than the
original machine that is showing the slow-down.
*I would assume, that your MPI bandwidth calculation only accounts for the
buffer size (i.e., only read or write).*
This is a good point, although I'm not sure why the bandwidth would be 1.5x
higher on the other machine that has very similar memory performance.
One more piece of information: The single-socket machine (Xeon E5-2650 v4)
has four RAM sticks in a quad-channel configuration, all tied to the same
socket as you can see in lstopo. On the dual-socket machine (the machine in
question), the four RAM sticks are in a dual-channel configuration, with
two sticks on each socket. So, I'm not sure if the dual- vs quad-channel
configuration is hurting maximum memory bandwidth per socket on the
dual-socket machine, despite the total bandwidths showing approximately the
same in STREAM.
I have tried to make the tests more uniform by binding the processes to the
same core. But that still calls into question the total memory bandwidth of
the dual-channel vs quad-channel memory configuration.
Do you have any thoughts on this? The question I'm trying to answer is: for
what reasons would the memory bandwidth be significantly lower on the
dual-socket Dell machine? I understand that making comparisons across
machines is tricky, but I've tried to provide as much information as
possible to isolate the key aspects of the memory configuration.
Thanks in advance!
Sam
On Thu, Aug 14, 2025 at 12:41 PM Boyle, Peter <pboyle(a)bnl.gov> wrote:
> Hi,
>
> 1)
> Obviously if there are multiple rank pairs on the node communicating
> concurrently, then multiple cores will be active copying concurrently
> increasing
> aggregate bandwidth to closer to the many core or threaded throughput,
> which is what you
> would get if you compile STREAM with "-fopenmp”.
>
> The Xeon part you quote has a peak of 76GB/s memory
> bandwidth per socket and multiple cores will be needed to saturate that.
>
> 2) It is common for people to use hybrid OpenMP and MPI, to minimize
> intra-node
> copy overhead. Often using one rank per NUMA domain.
> In that context, using MPI_Comm_split(…, MPI_COMM_TYPE_SHARED) will reveal
> which ranks can
> use unix shared memory regions (e.g. shmopen ) and then use ALL their
> threads concurrently
> to blast data between sockets, and use OpenMP within a socket.
>
> Then we get into the topic of careful NUMA binding of ranks to sockets
> etc...
>
> Best wishes,
>
> Peter
>
>
>
>
> *From: *Joachim Jenke via discuss <discuss(a)mpich.org>
> *Date: *Thursday, August 14, 2025 at 12:21 PM
> *To: *Sam Austin <sam.austin.p(a)gmail.com>
> *Cc: *Joachim Jenke <jenke(a)itc.rwth-aachen.de>, discuss(a)mpich.org <
> discuss(a)mpich.org>
> *Subject: *Re: [mpich-discuss] MPICH: SHM bandwidth very low on IPC test
>
> Hi Sam,
>
> the 10GB/s stream bandwidth calculation includes the number of
> read+written bytes (see lines 190/366).
>
> I would assume, that your MPI bandwidth calculation only accounts for
> the buffer size (i.e., only read or write). In shm communication one
> process (and therefore one core) streams/memcopies the data from the
> send to the receive buffer. So, when you see 3.5GB send bandwidth, that
> actually compares to 7GB of stream Copy bandwidth.
>
> As a side-effect of shm communication, we have actually seen that the
> placement of the copying process can determine the first-touch
> allocation of the buffer. Even if the memory is allocated with calloc,
> the memory is not paged. A bcast/scatter to node-local processes can
> result in paging all buffers to the same socket (what you typically want
> to avoid).
>
> Best
> Joachim
>
> Am 14.08.25 um 07:16 schrieb Sam Austin:
> > Hi Joachim,
> >
> > Thanks for this suggestion! I used stream to test the single-core memory
> > bandwidth. I am running on a Xeon E5-2699A v4, which has 55MB last level
> > cache. So, I ran with 30 million elements per the instructions. It
> > appears that I am seeing about 10 GB/s if I'm reading that right? If so,
> > I am still not sure why I am only seeing ~3.5 GB/s on shared memory
> > performance with MPICH.
> >
> > -------------------------------------------------------------
> > STREAM version $Revision: 5.10 $
> > -------------------------------------------------------------
> > This system uses 8 bytes per array element.
> > -------------------------------------------------------------
> > Array size = 30000000 (elements), Offset = 0 (elements)
> > Memory per array = 228.9 MiB (= 0.2 GiB).
> > Total memory required = 686.6 MiB (= 0.7 GiB).
> > Each kernel will be executed 10 times.
> > The *best* time for each kernel (excluding the first iteration)
> > will be used to compute the reported bandwidth.
> > -------------------------------------------------------------
> > Your clock granularity/precision appears to be 1 microseconds.
> > Each test below will take on the order of 30276 microseconds.
> > (= 30276 clock ticks)
> > Increase the size of the arrays if this shows that
> > you are not getting at least 20 clock ticks per test.
> > -------------------------------------------------------------
> > WARNING -- The above is only a rough guideline.
> > For best results, please be sure you know the
> > precision of your system timer.
> > -------------------------------------------------------------
> > Function Best Rate MB/s Avg time Min time Max time
> > Copy: 10038.1 0.048342 0.047818 0.050004
> > Scale: 10342.2 0.048738 0.046412 0.056605
> > Add: 10580.3 0.068542 0.068051 0.069805
> > Triad: 10703.0 0.067615 0.067271 0.068143
> > -------------------------------------------------------------
> > Solution Validates: avg error less than 1.000000e-13 on all three arrays
> > -------------------------------------------------------------
> >
> > Thanks,
> > Sam
> >
> > On Wed, Aug 13, 2025 at 5:26 PM Jenke, Joachim <jenke(a)itc.rwth-aachen.de
> > <mailto:[email protected] <jenke(a)itc.rwth-aachen.de>>> wrote:
> >
> > Hi Sam,
> >
> > Can you try out stream to understand the single-core memory
> > bandwidth of the system?
> >
> > https://urldefense.us/v3/__https://www.cs.virginia.edu/stream/ref.html__;!!… <https://
> > https://urldefense.us/v3/__http://www.cs.virginia.edu/stream/ref.html__;!!G… >
> <https://urldefense.us/v3/__http://www.cs.virginia.edu/stream/ref.html*3E__;… >
> >
> > Copy bandwidth for large junks (exceeding cache sizes) should
> > provide you an upper bound for shm communication bandwidth.
> >
> > Best
> > Joachim
> >
> > Am 13.08.2025 22:04 schrieb Sam Austin via discuss
> > <discuss(a)mpich.org <mailto:[email protected] <discuss(a)mpich.org>>>:
> > Hi all, I am working to configure MPICH and run a few examples on my
> > standalone server (single node). Here are the system specs: Server:
> > Dell PowerEdge C4130 CPUs: 2x Xeon E5-2699A v4 GPUs: 4x Tesla V100s
> > connected with NVLink, tied to motherboard
> > ZjQcmQRYFpfptBannerStart
> > This Message Is From an External Sender
> > This message came from outside your organization.
> > ZjQcmQRYFpfptBannerEnd
> > Hi all,
> >
> > I am working to configure MPICH and run a few examples on my
> > standalone server (single node). Here are the system specs:
> > Server: Dell PowerEdge C4130
> > CPUs: 2x Xeon E5-2699A v4
> > GPUs: 4x Tesla V100s connected with NVLink, tied to motherboard with
> > PCIe gen 3
> > OS: Ubuntu 24.04 LTS
> > I intend to use this system to develop multi-process programs for
> > eventual execution in a large, distributed HPC environment. I ran a
> > few tests with and without CUDA support; here is my mpichversion
> output:
> >
> > MPICH Version: 4.3.1
> > MPICH Release date: Fri Jun 20 09:24:41 AM CDT 2025
> > MPICH ABI: 17:1:5
> > MPICH Device: ch4:ofi
> > MPICH configure: --prefix=/opt/mpich/4.2.1-cpu --without-cuda
> > MPICH CC: gcc -O2
> > MPICH CXX: g++ -O2
> > MPICH F77: gfortran -O2
> > MPICH FC: gfortran -O2
> > MPICH features: threadcomm
> >
> > The first example that I ran was a bandwidth test for CPU-CPU and
> > GPU-GPU communication. This simple program sends small packets back
> > and forth between processes to test the bandwidth over the various
> > intra-node networks.
> >
> > The GPU-GPU bandwidth test showed that the GPU interconnect was
> > saturating at ~45 GB/s, which is nominal for the NVLink interconnect
> > topology present on the node (this was run with a CUDA-aware build
> > of MPICH). The problem appears during the CPU-CPU IPC test. In
> > theory, this test is pretty vanilla, as it is communicating between
> > processes using shared memory, and does not involve traversing any
> > of the intra-node networks (PCIe or NVLink). My understanding is
> > that the bandwidth observed on the CPU-CPU IPC test should be quite
> > high, at least higher than 10 GB/s.
> >
> > However, the intra-node IPC bandwidth appears to be very low, around
> > 3.5 GB/s, when running this test. I tried the following fixes in an
> > attempt to force MPICH to use shared memory, but to no avail:
> > Passing the option to explicitly specify `nemesis` during the build
> > configuration: "--with-device=ch3:nemesis --with-cuda"
> > Passing the option to explicitly specify shared memory with ch4 to
> > the configuration: "--with-ch4-shmmods=posix --with-cuda"
> > Rebuilding MPICH without GPU support: "--without-cuda"
> > Switching to Open MPI and running the same test
> > These results, especially the last one in which I saw the same
> > issues when running with Open MPI, makes me think it might be an
> > issue with my system configuration. The question is: why is the IPC
> > bandwidth so low despite supposedly using the SHM protocol? I'm
> > wondering if anyone has encountered this issue before or might be
> > able to lend some advice here. Any help would be greatly appreciated!
> >
> > Some interesting observations from the output below: when I run with
> > "mpiexec -np 2 -genv FI_PROVIDER=shm ...", the log file reports
> > "Opened fabric: shm". However, when I run without "-genv
> > FI_PROVIDER=shm", the log file reports "Opened fabric: 10.133.0.0/21
> > <https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!
> > ZaUD7Nw-
> > pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> > kgoP1Dp-C6$>", which I believe means that MPICH is falling back on
> > the TCP socket protocol. In this case, my key point of confusion is
> > that the observed bandwidth is essentially the same between the SHM
> > and TCP protocols. Perhaps my test script isn't set up properly?
> >
> > Thanks,
> > Sam
> >
> > The following is attached below:
> > Bandwidth test program
> > Run script for the program
> > Output of the script on my machine
> >
> ----------------------------------------------------------------------------------------------------------------
> > In case the attachment doesn't go through, here are the contents of
> > my test program, "shmem_check.cpp":
> >
> > // shmem_check.cpp
> > //
> > // This is a minimal benchmark to test the raw bandwidth of MPI
> > communication
> > // between two processes on the same node, using only host (CPU)
> memory.
> > // It completely removes CUDA to isolate the performance of the MPI
> > library's
> > // on-node communication mechanism (e.g., shared memory vs. TCP
> > loopback).
> > //
> > // Compile/run:
> > // /opt/mpich/4.2.1-cpu/bin/mpicxx -std=c++17 -I/opt/mpich/4.2.1-
> > cpu/include shmem_check.cpp -o shmem_check
> > // /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 ./shmem_check
> >
> > #include <iostream>
> > #include <vector>
> > #include <numeric>
> > #include <mpi.h>
> >
> > int main(int argc, char* argv[]) {
> > MPI_Init(&argc, &argv);
> >
> > int rank, size;
> > MPI_Comm_rank(MPI_COMM_WORLD, &rank);
> > MPI_Comm_size(MPI_COMM_WORLD, &size);
> >
> > if (size != 2) {
> > if (rank == 0) {
> > std::cerr << "Error: This program must be run with
> > exactly 2 MPI processes." << std::endl;
> > }
> > MPI_Finalize();
> > return 1;
> > }
> >
> > const int num_samples = 100;
> > const long long packet_size = 1LL << 28; // 256 MB
> >
> > // Allocate standard host memory. 'new' is sufficient.
> > char* buffer = new char[packet_size];
> >
> > if (rank == 0) {
> > std::cout << "--- Starting Host-to-Host MPI Bandwidth Test
> > ---" << std::endl;
> > std::cout << "Packet Size: " << (packet_size / (1024*1024))
> > << " MB" << std::endl;
> > }
> >
> > std::vector<double> timings;
> > for (int i = 0; i < num_samples; ++i) {
> > MPI_Barrier(MPI_COMM_WORLD);
> > double start_time = MPI_Wtime();
> >
> > if (rank == 0) {
> > MPI_Send(buffer, packet_size, MPI_CHAR, 1, 0,
> > MPI_COMM_WORLD);
> > MPI_Recv(buffer, 1, MPI_CHAR, 1, 1, MPI_COMM_WORLD,
> > MPI_STATUS_IGNORE); // Wait for confirmation
> > } else { // rank == 1
> > MPI_Recv(buffer, packet_size, MPI_CHAR, 0, 0,
> > MPI_COMM_WORLD, MPI_STATUS_IGNORE);
> > MPI_Send(buffer, 1, MPI_CHAR, 0, 1, MPI_COMM_WORLD); //
> > Send confirmation
> > }
> >
> > double end_time = MPI_Wtime();
> > if (i >= 10) { // Discard warmup runs
> > timings.push_back(end_time - start_time);
> > }
> > }
> >
> > if (rank == 0) {
> > double total_time = std::accumulate(timings.begin(),
> > timings.end(), 0.0);
> > double avg_time = total_time / timings.size();
> > double bandwidth = (static_cast<double>(packet_size) /
> > (1024.0 * 1024.0 * 1024.0)) / avg_time;
> >
> > std::cout <<
> > "------------------------------------------------" << std::endl;
> > std::cout << "Average Host-to-Host Bandwidth: " <<
> > bandwidth << " GB/s" << std::endl;
> > std::cout <<
> > "------------------------------------------------" << std::endl;
> > }
> >
> > // Clean up host memory
> > delete[] buffer;
> >
> > MPI_Finalize();
> > return 0;
> > }
> >
> >
> ----------------------------------------------------------------------------------------------------------------
> > Here is the script to run the test with verbose compilation and the
> > `shm` layer forced and unforced:
> >
> > #!/usr/bin/zsh
> > source ~/.zshrc
> >
> > # Compile
> > /opt/mpich/4.2.1-cpu/bin/mpicxx -std=c++17 -I/opt/mpich/4.2.1-cpu/
> > include shmem_check.cpp -o shmem_check
> >
> > # Run with shm forced
> > /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 -genv FI_PROVIDER=shm -genv
> > FI_LOG_LEVEL=debug ./shmem_check 2> output_shm.txt
> >
> > # Run without shm forced
> > /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 -genv FI_LOG_LEVEL=debug ./
> > shmem_check 2> output_no_shm.txt
> >
> > echo "Output of script with SHM forced: "
> > grep -i "opened fabric" output_shm.txt
> >
> > echo "Output of script with SHM not forced: "
> > grep -i "opened fabric" output_no_shm.txt
> >
> >
> ----------------------------------------------------------------------------------------------------------------
> > Here is the output :
> >
> > --- Starting Host-to-Host MPI Bandwidth Test ---
> > Packet Size: 256 MB
> > ------------------------------------------------
> > Average Host-to-Host Bandwidth: 3.35709 GB/s
> > ------------------------------------------------
> > --- Starting Host-to-Host MPI Bandwidth Test ---
> > Packet Size: 256 MB
> > ------------------------------------------------
> > Average Host-to-Host Bandwidth: 3.54924 GB/s
> > ------------------------------------------------
> > Output of script with SHM forced:
> > libfabric:3174297:1755114546::core:core:fi_fabric_():1503<info>
> > Opened fabric: shm
> > libfabric:3174298:1755114546::core:core:fi_fabric_():1503<info>
> > Opened fabric: shm
> > Output of script with SHM not forced:
> > libfabric:3174351:1755114554::core:core:fi_fabric_():1503<info>
> > Opened fabric: 10.133.0.0/21 <https://urldefense.us/v3/
> > __https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__;!!G_uCfscf7eWS!aD-lRdCn7hGKD7zH2zy7Xu9dMPxA6Ww49nwtLe35Q6nnbgIDUcWcv5UzH-9BEb3gC0GN43yKL5mkCcxTEnIl$
> > pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> > kgoP1Dp-C6$>
> > libfabric:3174350:1755114554::core:core:fi_fabric_():1503<info>
> > Opened fabric: 10.133.0.0/21 <https://urldefense.us/v3/
> > __https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__;!!G_uCfscf7eWS!aD-lRdCn7hGKD7zH2zy7Xu9dMPxA6Ww49nwtLe35Q6nnbgIDUcWcv5UzH-9BEb3gC0GN43yKL5mkCcxTEnIl$
> > pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> > kgoP1Dp-C6$>
> >
>
>
> --
> Dr. rer. nat. Joachim Jenke
> Deputy Group Lead
>
> IT Center
> Group: HPC - Parallelism, Runtime Analysis & Machine Learning
> Division: Computational Science and Engineering
> RWTH Aachen University
> Seffenter Weg 23
> D 52074 Aachen (Germany)
> Tel: +49 241 80- 24765
> Fax: +49 241 80-624765
> jenke(a)itc.rwth-aachen.de
> https://urldefense.us/v3/__http://www.itc.rwth-aachen.de__;!!G_uCfscf7eWS!a…
>
1
0
Hi,
1)
Obviously if there are multiple rank pairs on the node communicating
concurrently, then multiple cores will be active copying concurrently increasing
aggregate bandwidth to closer to the many core or threaded throughput, which is what you
would get if you compile STREAM with "-fopenmp”.
The Xeon part you quote has a peak of 76GB/s memory
bandwidth per socket and multiple cores will be needed to saturate that.
2) It is common for people to use hybrid OpenMP and MPI, to minimize intra-node
copy overhead. Often using one rank per NUMA domain.
In that context, using MPI_Comm_split(…, MPI_COMM_TYPE_SHARED) will reveal which ranks can
use unix shared memory regions (e.g. shmopen ) and then use ALL their threads concurrently
to blast data between sockets, and use OpenMP within a socket.
Then we get into the topic of careful NUMA binding of ranks to sockets etc...
Best wishes,
Peter
From: Joachim Jenke via discuss <discuss(a)mpich.org>
Date: Thursday, August 14, 2025 at 12:21 PM
To: Sam Austin <sam.austin.p(a)gmail.com>
Cc: Joachim Jenke <jenke(a)itc.rwth-aachen.de>, discuss(a)mpich.org <discuss(a)mpich.org>
Subject: Re: [mpich-discuss] MPICH: SHM bandwidth very low on IPC test
Hi Sam,
the 10GB/s stream bandwidth calculation includes the number of
read+written bytes (see lines 190/366).
I would assume, that your MPI bandwidth calculation only accounts for
the buffer size (i.e., only read or write). In shm communication one
process (and therefore one core) streams/memcopies the data from the
send to the receive buffer. So, when you see 3.5GB send bandwidth, that
actually compares to 7GB of stream Copy bandwidth.
As a side-effect of shm communication, we have actually seen that the
placement of the copying process can determine the first-touch
allocation of the buffer. Even if the memory is allocated with calloc,
the memory is not paged. A bcast/scatter to node-local processes can
result in paging all buffers to the same socket (what you typically want
to avoid).
Best
Joachim
Am 14.08.25 um 07:16 schrieb Sam Austin:
> Hi Joachim,
>
> Thanks for this suggestion! I used stream to test the single-core memory
> bandwidth. I am running on a Xeon E5-2699A v4, which has 55MB last level
> cache. So, I ran with 30 million elements per the instructions. It
> appears that I am seeing about 10 GB/s if I'm reading that right? If so,
> I am still not sure why I am only seeing ~3.5 GB/s on shared memory
> performance with MPICH.
>
> -------------------------------------------------------------
> STREAM version $Revision: 5.10 $
> -------------------------------------------------------------
> This system uses 8 bytes per array element.
> -------------------------------------------------------------
> Array size = 30000000 (elements), Offset = 0 (elements)
> Memory per array = 228.9 MiB (= 0.2 GiB).
> Total memory required = 686.6 MiB (= 0.7 GiB).
> Each kernel will be executed 10 times.
> The *best* time for each kernel (excluding the first iteration)
> will be used to compute the reported bandwidth.
> -------------------------------------------------------------
> Your clock granularity/precision appears to be 1 microseconds.
> Each test below will take on the order of 30276 microseconds.
> (= 30276 clock ticks)
> Increase the size of the arrays if this shows that
> you are not getting at least 20 clock ticks per test.
> -------------------------------------------------------------
> WARNING -- The above is only a rough guideline.
> For best results, please be sure you know the
> precision of your system timer.
> -------------------------------------------------------------
> Function Best Rate MB/s Avg time Min time Max time
> Copy: 10038.1 0.048342 0.047818 0.050004
> Scale: 10342.2 0.048738 0.046412 0.056605
> Add: 10580.3 0.068542 0.068051 0.069805
> Triad: 10703.0 0.067615 0.067271 0.068143
> -------------------------------------------------------------
> Solution Validates: avg error less than 1.000000e-13 on all three arrays
> -------------------------------------------------------------
>
> Thanks,
> Sam
>
> On Wed, Aug 13, 2025 at 5:26 PM Jenke, Joachim <jenke(a)itc.rwth-aachen.de
> <mailto:[email protected]>> wrote:
>
> Hi Sam,
>
> Can you try out stream to understand the single-core memory
> bandwidth of the system?
>
> https://urldefense.us/v3/__https://www.cs.virginia.edu/stream/ref.html__;!!… <https://
> https://urldefense.us/v3/__http://www.cs.virginia.edu/stream/ref.html__;!!G… ><https://urldefense.us/v3/__http://www.cs.virginia.edu/stream/ref.html__;!!G… >>
>
> Copy bandwidth for large junks (exceeding cache sizes) should
> provide you an upper bound for shm communication bandwidth.
>
> Best
> Joachim
>
> Am 13.08.2025 22:04 schrieb Sam Austin via discuss
> <discuss(a)mpich.org <mailto:[email protected]>>:
> Hi all, I am working to configure MPICH and run a few examples on my
> standalone server (single node). Here are the system specs: Server:
> Dell PowerEdge C4130 CPUs: 2x Xeon E5-2699A v4 GPUs: 4x Tesla V100s
> connected with NVLink, tied to motherboard
> ZjQcmQRYFpfptBannerStart
> This Message Is From an External Sender
> This message came from outside your organization.
> ZjQcmQRYFpfptBannerEnd
> Hi all,
>
> I am working to configure MPICH and run a few examples on my
> standalone server (single node). Here are the system specs:
> Server: Dell PowerEdge C4130
> CPUs: 2x Xeon E5-2699A v4
> GPUs: 4x Tesla V100s connected with NVLink, tied to motherboard with
> PCIe gen 3
> OS: Ubuntu 24.04 LTS
> I intend to use this system to develop multi-process programs for
> eventual execution in a large, distributed HPC environment. I ran a
> few tests with and without CUDA support; here is my mpichversion output:
>
> MPICH Version: 4.3.1
> MPICH Release date: Fri Jun 20 09:24:41 AM CDT 2025
> MPICH ABI: 17:1:5
> MPICH Device: ch4:ofi
> MPICH configure: --prefix=/opt/mpich/4.2.1-cpu --without-cuda
> MPICH CC: gcc -O2
> MPICH CXX: g++ -O2
> MPICH F77: gfortran -O2
> MPICH FC: gfortran -O2
> MPICH features: threadcomm
>
> The first example that I ran was a bandwidth test for CPU-CPU and
> GPU-GPU communication. This simple program sends small packets back
> and forth between processes to test the bandwidth over the various
> intra-node networks.
>
> The GPU-GPU bandwidth test showed that the GPU interconnect was
> saturating at ~45 GB/s, which is nominal for the NVLink interconnect
> topology present on the node (this was run with a CUDA-aware build
> of MPICH). The problem appears during the CPU-CPU IPC test. In
> theory, this test is pretty vanilla, as it is communicating between
> processes using shared memory, and does not involve traversing any
> of the intra-node networks (PCIe or NVLink). My understanding is
> that the bandwidth observed on the CPU-CPU IPC test should be quite
> high, at least higher than 10 GB/s.
>
> However, the intra-node IPC bandwidth appears to be very low, around
> 3.5 GB/s, when running this test. I tried the following fixes in an
> attempt to force MPICH to use shared memory, but to no avail:
> Passing the option to explicitly specify `nemesis` during the build
> configuration: "--with-device=ch3:nemesis --with-cuda"
> Passing the option to explicitly specify shared memory with ch4 to
> the configuration: "--with-ch4-shmmods=posix --with-cuda"
> Rebuilding MPICH without GPU support: "--without-cuda"
> Switching to Open MPI and running the same test
> These results, especially the last one in which I saw the same
> issues when running with Open MPI, makes me think it might be an
> issue with my system configuration. The question is: why is the IPC
> bandwidth so low despite supposedly using the SHM protocol? I'm
> wondering if anyone has encountered this issue before or might be
> able to lend some advice here. Any help would be greatly appreciated!
>
> Some interesting observations from the output below: when I run with
> "mpiexec -np 2 -genv FI_PROVIDER=shm ...", the log file reports
> "Opened fabric: shm". However, when I run without "-genv
> FI_PROVIDER=shm", the log file reports "Opened fabric: 10.133.0.0/21
> <https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!
> ZaUD7Nw-
> pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> kgoP1Dp-C6$>", which I believe means that MPICH is falling back on
> the TCP socket protocol. In this case, my key point of confusion is
> that the observed bandwidth is essentially the same between the SHM
> and TCP protocols. Perhaps my test script isn't set up properly?
>
> Thanks,
> Sam
>
> The following is attached below:
> Bandwidth test program
> Run script for the program
> Output of the script on my machine
> ----------------------------------------------------------------------------------------------------------------
> In case the attachment doesn't go through, here are the contents of
> my test program, "shmem_check.cpp":
>
> // shmem_check.cpp
> //
> // This is a minimal benchmark to test the raw bandwidth of MPI
> communication
> // between two processes on the same node, using only host (CPU) memory.
> // It completely removes CUDA to isolate the performance of the MPI
> library's
> // on-node communication mechanism (e.g., shared memory vs. TCP
> loopback).
> //
> // Compile/run:
> // /opt/mpich/4.2.1-cpu/bin/mpicxx -std=c++17 -I/opt/mpich/4.2.1-
> cpu/include shmem_check.cpp -o shmem_check
> // /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 ./shmem_check
>
> #include <iostream>
> #include <vector>
> #include <numeric>
> #include <mpi.h>
>
> int main(int argc, char* argv[]) {
> MPI_Init(&argc, &argv);
>
> int rank, size;
> MPI_Comm_rank(MPI_COMM_WORLD, &rank);
> MPI_Comm_size(MPI_COMM_WORLD, &size);
>
> if (size != 2) {
> if (rank == 0) {
> std::cerr << "Error: This program must be run with
> exactly 2 MPI processes." << std::endl;
> }
> MPI_Finalize();
> return 1;
> }
>
> const int num_samples = 100;
> const long long packet_size = 1LL << 28; // 256 MB
>
> // Allocate standard host memory. 'new' is sufficient.
> char* buffer = new char[packet_size];
>
> if (rank == 0) {
> std::cout << "--- Starting Host-to-Host MPI Bandwidth Test
> ---" << std::endl;
> std::cout << "Packet Size: " << (packet_size / (1024*1024))
> << " MB" << std::endl;
> }
>
> std::vector<double> timings;
> for (int i = 0; i < num_samples; ++i) {
> MPI_Barrier(MPI_COMM_WORLD);
> double start_time = MPI_Wtime();
>
> if (rank == 0) {
> MPI_Send(buffer, packet_size, MPI_CHAR, 1, 0,
> MPI_COMM_WORLD);
> MPI_Recv(buffer, 1, MPI_CHAR, 1, 1, MPI_COMM_WORLD,
> MPI_STATUS_IGNORE); // Wait for confirmation
> } else { // rank == 1
> MPI_Recv(buffer, packet_size, MPI_CHAR, 0, 0,
> MPI_COMM_WORLD, MPI_STATUS_IGNORE);
> MPI_Send(buffer, 1, MPI_CHAR, 0, 1, MPI_COMM_WORLD); //
> Send confirmation
> }
>
> double end_time = MPI_Wtime();
> if (i >= 10) { // Discard warmup runs
> timings.push_back(end_time - start_time);
> }
> }
>
> if (rank == 0) {
> double total_time = std::accumulate(timings.begin(),
> timings.end(), 0.0);
> double avg_time = total_time / timings.size();
> double bandwidth = (static_cast<double>(packet_size) /
> (1024.0 * 1024.0 * 1024.0)) / avg_time;
>
> std::cout <<
> "------------------------------------------------" << std::endl;
> std::cout << "Average Host-to-Host Bandwidth: " <<
> bandwidth << " GB/s" << std::endl;
> std::cout <<
> "------------------------------------------------" << std::endl;
> }
>
> // Clean up host memory
> delete[] buffer;
>
> MPI_Finalize();
> return 0;
> }
>
> ----------------------------------------------------------------------------------------------------------------
> Here is the script to run the test with verbose compilation and the
> `shm` layer forced and unforced:
>
> #!/usr/bin/zsh
> source ~/.zshrc
>
> # Compile
> /opt/mpich/4.2.1-cpu/bin/mpicxx -std=c++17 -I/opt/mpich/4.2.1-cpu/
> include shmem_check.cpp -o shmem_check
>
> # Run with shm forced
> /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 -genv FI_PROVIDER=shm -genv
> FI_LOG_LEVEL=debug ./shmem_check 2> output_shm.txt
>
> # Run without shm forced
> /opt/mpich/4.2.1-cpu/bin/mpiexec -np 2 -genv FI_LOG_LEVEL=debug ./
> shmem_check 2> output_no_shm.txt
>
> echo "Output of script with SHM forced: "
> grep -i "opened fabric" output_shm.txt
>
> echo "Output of script with SHM not forced: "
> grep -i "opened fabric" output_no_shm.txt
>
> ----------------------------------------------------------------------------------------------------------------
> Here is the output :
>
> --- Starting Host-to-Host MPI Bandwidth Test ---
> Packet Size: 256 MB
> ------------------------------------------------
> Average Host-to-Host Bandwidth: 3.35709 GB/s
> ------------------------------------------------
> --- Starting Host-to-Host MPI Bandwidth Test ---
> Packet Size: 256 MB
> ------------------------------------------------
> Average Host-to-Host Bandwidth: 3.54924 GB/s
> ------------------------------------------------
> Output of script with SHM forced:
> libfabric:3174297:1755114546::core:core:fi_fabric_():1503<info>
> Opened fabric: shm
> libfabric:3174298:1755114546::core:core:fi_fabric_():1503<info>
> Opened fabric: shm
> Output of script with SHM not forced:
> libfabric:3174351:1755114554::core:core:fi_fabric_():1503<info>
> Opened fabric: 10.133.0.0/21 <https://urldefense.us/v3/
> __https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__;!!G_uCfscf7eWS!c899LMcJ4ey07LaDwvVF8ATRKCNqV6XrOuml2yi_MgGn1lJPtYKvEC-o9Rx_pi5LmumFFPabsm7tCA$
> pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> kgoP1Dp-C6$>
> libfabric:3174350:1755114554::core:core:fi_fabric_():1503<info>
> Opened fabric: 10.133.0.0/21 <https://urldefense.us/v3/
> __https://urldefense.us/v3/__http://10.133.0.0/21__;!!G_uCfscf7eWS!ZaUD7Nw-__;!!G_uCfscf7eWS!c899LMcJ4ey07LaDwvVF8ATRKCNqV6XrOuml2yi_MgGn1lJPtYKvEC-o9Rx_pi5LmumFFPabsm7tCA$
> pSrvdcr4vb0JBtm7m5HhtE6d7G1wb5HakwqLQQenlo0WTl1tkzV3CrJnLwCQ7cVvC-
> kgoP1Dp-C6$>
>
--
Dr. rer. nat. Joachim Jenke
Deputy Group Lead
IT Center
Group: HPC - Parallelism, Runtime Analysis & Machine Learning
Division: Computational Science and Engineering
RWTH Aachen University
Seffenter Weg 23
D 52074 Aachen (Germany)
Tel: +49 241 80- 24765
Fax: +49 241 80-624765
jenke(a)itc.rwth-aachen.de
https://urldefense.us/v3/__http://www.itc.rwth-aachen.de__;!!G_uCfscf7eWS!c… <https://urldefense.us/v3/__http://www.itc.rwth-aachen.de__;!!G_uCfscf7eWS!c… >
1
0