discuss
Threads by month
- ----- 2026 -----
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
June 2026
- 3 participants
- 30 discussions
as less latency than 4KB.
I was looking for explanation of this behavior but did not get any.
1. MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE is set to 128KB. So none of the abov=
e message size is using Rendezvous protocol. Is there any partition inside =
eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB)? If yes then wh=
at are the boundaries for them? Can I log them with debug-event-logging?
Setup I am using:
- two nodes has intel core i7, one with 16gb memory another one 8gb
- mpich 3.2.1, configured and build to use nemesis tcp
- 1gb Ethernet connection
- NFS is using for sharing
- osu_latency: uses MPI_Send and MPI_Recv
- MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE=3D 131072 (128KB)
Can anyone help me on that? Thanks in advance.
Best Regards,
Abu Naser
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--
Jeff Hammond
jeff.science(a)gmail.com<mailto: jeff.science(a)gmail.com>
http: //jeffhammond.github.io/
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--_000_BLUPR0501MB2003D13F739833B52E59DBB097430BLUPR0501MB2003_
Content-Type: text/html; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
<html><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3DWindows-1=
252">
<style type=3D"text/css" style=3D"display: none;"><!-- P {margin-top:0;margi=
n-bottom: 0;} --></style>
</head>
<body dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt;color:#000000;font=
-family: Calibri,Helvetica,sans-serif;" dir=3D"ltr">
<p style=3D"margin-top: 0;margin-bottom:0">Hello Min,</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">Now for some cases it is working =
and for some cases not.</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"><b>Cases when it worked:</b></p>
<p style=3D"margin-top: 0;margin-bottom:0">- when application binary (e.g cp=
i, osu_bw, osu_latency) is compiled with other mpicc (generated when config=
ured with tcp), then mpiexec (<span>generated when</span> configured with s=
ock) could run it.</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"><b>Cases not working:</b></p>
<p style=3D"margin-top: 0;margin-bottom:0">- application binary (e.g cpi) is=
compiled with mpicc (<span>generated when</span> configured with sock), th=
en mpiexec (<span>generated when</span> configured with sock) could not run=
it and produce the same error message.
[lib path was set]</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">Thank you.<br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<div id=3D"Signature">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr style=3D"display: inline-block;width:98%" tabindex=3D"-1">
<div id=3D"divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif" st=
yle=3D"font-size: 11pt" color=3D"#000000"><b>From:</b> Min Si <msi(a)anl.go=
v><br>
<b>Sent: </b> Monday, July 2, 2018 2:10:23 PM<br>
<b>To: </b> discuss(a)mpich.org<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Could you please try mpich-3.3b3 ?<=
br>
<a class=3D"x_moz-txt-link-freetext" href=3D"http: //www.mpich.org/static/do=
wnloads/3.3b3/mpich-3.3b3.tar.gz">http: //www.mpich.org/static/downloads/3.3=
b3/mpich-3.3b3.tar.gz</a><br>
<br>
Min<br>
<div class=3D"x_moz-cite-prefix">On 2018/07/02 13: 01, Abu Naser wrote:<br>
</div>
<blockquote type=3D"cite"><style type=3D"text/css" style=3D"display: none">
<!--
p
{margin-top: 0;
margin-bottom:0}
-->
</style>
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p style=3D"margin-top: 0; margin-bottom:0">Hello Min,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I have downloaded it from <a hre=
f=3D"http: //www.mpich.org/static/downloads/3.2.1/mpich-3.2.1.tar.gz" class=
=3D"x_OWAAutoLink" id=3D"LPlnk943697">
http: //www.mpich.org/static/downloads/3.2.1/mpich-3.2.1.tar.gz</a> but it d=
id not work. I have received almost same error. Except this time no process=
information from my remote machine.</p>
<p style=3D"margin-top: 0; margin-bottom:0"><b>Previously I have received th=
is - </b>
<br>
</p>
<div style=3D""><i><span style=3D"font-size: 10pt; color:rgb(255,0,0)">Proce=
ss 3 of 4 is on dhcp16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt; color:rgb(255,0,0)">Proce=
ss 1 of 4 is on dhcp16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt; color:rgb(255,0,0)">Proce=
ss 0 of 4 is on dhcp16198</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt; color:rgb(255,0,0)">Proce=
ss 2 of 4 is on dhcp16198</span></i></div>
<p style=3D"margin-top: 0; margin-bottom:0"><b>With the new source code -</b=
></p>
<div><i><span style=3D"font-size: 10pt; color:rgb(255,0,0)">Process 0 of 4 i=
s on dhcp16198</span></i></div>
<div><i><span style=3D"font-size: 10pt; color:rgb(255,0,0)">Process 2 of 4 i=
s on dhcp16198</span></i></div>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><b>Entire error message is:</b><=
/p>
<div><i><span style=3D"font-size: 10pt">Process 0 of 4 is on dhcp16198</span=
></i></div>
<div><i><span style=3D"font-size: 10pt">Process 2 of 4 is on dhcp16198</span=
></i></div>
<div><i><span style=3D"font-size: 10pt">Fatal error in PMPI_Bcast: Unknown e=
rror class, error stack: </span></i></div>
<div><i><span style=3D"font-size: 10pt">PMPI_Bcast(1600)....................=
........: MPI_Bcast(buf=3D0x7ffd1ee145f0, count=3D1, MPI_INT, root=3D0, MPI=
_COMM_WORLD) failed</span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast_impl(1452)...............=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast(1476)....................=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast_intra(1249)..............=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1081)................=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(285)............=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIC_Send(303)......................=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIC_Wait(226)......................=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIDI_CH3i_Progress_wait(242).......=
........: an error occurred while handling an event returned by MPIDU_Sock_=
Wait()</span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIDI_CH3I_Progress_handle_sock_even=
t(698)..: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIDI_CH3_Sockconn_handle_connect_ev=
ent(597): [ch3:sock] failed to connnect to remote process</span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIDU_Socki_handle_connect(808).....=
........: connection failure (set=3D0,sock=3D1,errno=3D111:Connection refus=
ed)</span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1088)................=
........: </span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(310)............=
........: Failure during collective</span></i></div>
<div><i><span style=3D"font-size: 10pt">Fatal error in PMPI_Bcast: Other MPI=
error, error stack:</span></i></div>
<div><i><span style=3D"font-size: 10pt">PMPI_Bcast(1600)........: MPI_Bcast(=
buf=3D0x7ffe2eeb90f0, count=3D1, MPI_INT, root=3D0, MPI_COMM_WORLD) failed<=
/span></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast_impl(1452)...: </spa=
n></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast(1476)........: </spa=
n></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast_intra(1249)..: </spa=
n></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1088)....: </spa=
n></i></div>
<div><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(310): Failure du=
ring collective</span></i></div>
<br>
<p style=3D"margin-top: 0; margin-bottom:0">Again if I configure the new sou=
rce with tcp, it works fine.</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Thank You.<br>
</p>
<div id=3D"x_Signature">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1" style=3D"display: inline-block; width:98%">
<div id=3D"x_divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif" =
color=3D"#000000" style=3D"font-size: 11pt"><b>From:</b> Min Si
<a class=3D"x_moz-txt-link-rfc2396E" href=3D"mailto: msi(a)anl.gov"><msi@an=
l.gov></a><br>
<b>Sent: </b> Monday, July 2, 2018 11:56:51 AM<br>
<b>To: </b> <a class=3D"x_moz-txt-link-abbreviated" href=3D"mailto:discuss@m=
pich.org">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Hi Abu,<br>
<br>
Thanks for reporting this. Can you please try the latest release with ch3/s=
ock and see if you still have this error ?
<br>
<br>
Min<br>
<div class=3D"x_x_moz-cite-prefix">On 2018/07/01 21: 47, Abu Naser wrote:<br=
>
</div>
<blockquote type=3D"cite">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p style=3D"">Hello Min,</p>
<p style=3D""><br>
</p>
<p style=3D"">After compiling my mpich-3.2.1 with sock, while I was trying =
to run any program including osu benchmark or examples/cpi&=
nbsp; in two machines, I have received following error -</p>
<p style=3D""><br>
</p>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 3 of 4 is on dhcp=
16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 1 of 4 is on dhcp=
16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 0 of 4 is on dhcp=
16198</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 2 of 4 is on dhcp=
16198</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Fatal error in PMPI_Bcast=
: Unknown error class, error stack:</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">PMPI_Bcast(1600).........=
...................: MPI_Bcast(buf=3D0x7ffc1808542c, count=3D1, MPI_INT, ro=
ot=3D0, MPI_COMM_WORLD) failed</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_impl(1452)....=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast(1476).........=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_intra(1249)...=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1081).....=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(285).=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIC_Send(303)...........=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIC_Wait(226)...........=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDI_CH3i_Progress_wait(=
242)...............: an error occurred while handling an event returned by =
MPIDU_Sock_Wait()</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDI_CH3I_Progress_handl=
e_sock_event(698)..: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDI_CH3_Sockconn_handle=
_connect_event(597): [ch3:sock] failed to connnect to remote process</span>=
</i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDU_Socki_handle_connec=
t(808).............: connection failure (set=3D0,sock=3D1,errno=3D111:Conne=
ction refused)</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1088).....=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(310).=
...................: Failure during collective</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Fatal error in PMPI_Bcast=
: Other MPI error, error stack:</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">PMPI_Bcast(1600)........:=
MPI_Bcast(buf=3D0x7ffd9eeebdac, count=3D1, MPI_INT, root=3D0, MPI_COMM_WOR=
LD) failed</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_impl(1452)...:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast(1476)........:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_intra(1249)..:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1088)....:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(310):=
Failure during collective</span></i></div>
<br style=3D"">
<p style=3D""><span style=3D"font-size: 12pt">I checked the mpich FAQ a=
nd also mpich discussion list. Based on that I have checked </span>fol=
lowings<span style=3D"font-size: 12pt"> </span><span style=3D"font-size=
: 12pt">and found they are fine in my machines -</span><br>
</p>
<p style=3D""><span style=3D"font-size: 12pt">- firewall is disabled in both=
machine</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- I can do </span>passwor=
d less<span style=3D"font-size: 12pt"> ssh in both machine</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- /etc/hosts in both machine c=
onfigured with ip address and name properly</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- I have updated the library p=
ath and used absolute path for mpiexec</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- Most importantly when I conf=
igured and build mpich with tcp, it works fine.</span></p>
<p style=3D""><span style=3D"font-size: 12pt"><br>
</span></p>
<p style=3D""><span style=3D"font-size: 12pt"> I think I am </span=
><span style=3D"font-size: 12pt">missing something but could not figure out =
yet. Any help would be
</span>appreciated<span style=3D"font-size: 12pt">.</span></p>
<p style=3D""><span style=3D"font-size: 12pt"><br>
</span></p>
<p style=3D""><span style=3D"font-size: 12pt">Thank you.</span></p>
<br>
<p style=3D""><br>
</p>
<p style=3D""><br>
</p>
<p style=3D""><br>
</p>
<div id=3D"x_x_Signature" style=3D"">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1" style=3D"display: inline-block; width:98%">
<div id=3D"x_x_divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif=
" color=3D"#000000" style=3D"font-size: 11pt"><b>From:</b> Min Si
<a class=3D"x_x_moz-txt-link-rfc2396E x_OWAAutoLink" href=3D"mailto: msi@anl=
.gov" id=3D"LPlnk191149">
<msi(a)anl.gov></a><br>
<b>Sent: </b> Tuesday, June 26, 2018 12:54:29 PM<br>
<b>To: </b> <a class=3D"x_x_moz-txt-link-abbreviated x_OWAAutoLink" href=3D"=
mailto: discuss(a)mpich.org" id=3D"LPlnk414203">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Hi Abu,<br>
<br>
I think the results are stable enough. Perhaps you could also try the follo=
wing tests, and see if similar trend exists: <br>
- MPICH/socket (set `--with-device=3Dch3: sock` at configure)<br>
- A socket-based pingpong test without MPI. <br>
<br>
At this point, I could not think of any MPI-specific design for 2k/8k messa=
ges. My guess is that it is related to your network connection.
<br>
<br>
Min<br>
<br>
<div class=3D"x_x_x_moz-cite-prefix">On 2018/06/24 11: 09, Abu Naser wrote:<=
br>
</div>
<blockquote type=3D"cite">
<meta content=3D"text/html; charset=3DWindows-1252">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Min and Jeff,</p>
<p><br>
</p>
<p>Here is my experiment results. Default number of iterations in osu_=
latency for 0B =96 8KB is 10,000. With that setting I had run the osu_laten=
cy 100 times and found standard deviation 33 for 8KB message size.</p>
<p><br>
</p>
<p>So later I have set the iteration to 50,000 and 100,000 for 1KB =96 16KB=
message size. Then run osu_latency for 100 times for each setting and take=
the average and standard deviation.</p>
<p><br>
</p>
<table width=3D"665">
<colgroup><col width=3D"99"><col width=3D"112"><col width=3D"118"><col widt=
h=3D"154"><col width=3D"140"></colgroup>
<tbody>
<tr>
<td width=3D"99">
<p><b>Msg Size in Bytes</b></p>
</td>
<td width=3D"112">
<p><b>Avg time in us (50K iterations)</b></p>
</td>
<td width=3D"118">
<p><b>Avg time in us (100k iterations)</b></p>
</td>
<td width=3D"154">
<p><b>Standard deviation (50K iterations)</b></p>
</td>
<td width=3D"140">
<p><b>Standard deviation (100K iterations)</b></p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>1k</p>
</td>
<td width=3D"112">
<p>85.10</p>
</td>
<td width=3D"118">
<p>84.9</p>
</td>
<td width=3D"154">
<p>0.55</p>
</td>
<td width=3D"140">
<p>0.45</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>2k</p>
</td>
<td width=3D"112">
<p>75.79</p>
</td>
<td width=3D"118">
<p>74.63</p>
</td>
<td width=3D"154">
<p>5.09</p>
</td>
<td width=3D"140">
<p>4.44</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>4k</p>
</td>
<td width=3D"112">
<p>273.80</p>
</td>
<td width=3D"118">
<p>274.71</p>
</td>
<td width=3D"154">
<p>4.18</p>
</td>
<td width=3D"140">
<p>2.45</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>8k</p>
</td>
<td width=3D"112">
<p>258.56</p>
</td>
<td width=3D"118">
<p>249.83</p>
</td>
<td width=3D"154">
<p>21.14</p>
</td>
<td width=3D"140">
<p>28</p>
</td>
</tr>
<tr>
<td height=3D"24" width=3D"99">
<p>16k</p>
</td>
<td width=3D"112">
<p>281.31</p>
</td>
<td width=3D"118">
<p>281.02</p>
</td>
<td width=3D"154">
<p>3.22</p>
</td>
<td width=3D"140">
<p>4.10</p>
</td>
</tr>
</tbody>
</table>
<p><br>
</p>
<p><br>
</p>
<p>The standard deviation of 8K message is so high and that implies it actu=
ally not producing any consistent latency time. Looks like that's the =
reason for 8K is taking less time than 4K.</p>
<p><br>
</p>
<p>Meanwhile, 2K has standard deviation less than 5 but 1K message latency =
timing are more densely populated than 2K. So probably this is the explanat=
ion for 2K message less latency time.</p>
<p><br>
</p>
<p>Thank you for your suggestions.</p>
<br>
<p><br>
</p>
<div id=3D"x_x_x_Signature">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Abu Naser<br>
<b>Sent: </b> Wednesday, June 20, 2018 1:48:53 PM<br>
<b>To: </b> <a class=3D"x_x_x_moz-txt-link-abbreviated x_x_OWAAutoLink" href=
=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk729146">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3Diso-8859-1">
<div dir=3D"ltr">
<div id=3D"x_x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<div id=3D"x_x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Min,</p>
<p><br>
</p>
<p>Thanks for the clarification. I will do the experiment.<br>
</p>
<p><br>
</p>
<div id=3D"x_x_x_x_Signature">
<div id=3D"x_x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Thanks.</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Min Si <a class=
=3D"x_x_x_moz-txt-link-rfc2396E x_x_OWAAutoLink" href=3D"mailto: msi(a)anl.gov=
" id=3D"LPlnk558260">
<msi(a)anl.gov></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 1:39:30 PM<br>
<b>To: </b> <a class=3D"x_x_x_moz-txt-link-abbreviated x_x_OWAAutoLink" href=
=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk472728">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div>Hi Abu,<br>
<br>
I think Jeff means that you should run your experiment with more iterations=
in order to get a stable results.<br>
- Increase the iteration of for loop in each execution (I think osu benchma=
rk allows you to set it)<br>
- Run the experiments 10 or 100 times, and take the average and standard de=
viation.<br>
<br>
If you see a very small standard deviation (e.g., <=3D5%), then the tren=
d is stable and you might not see such gaps.<br>
<br>
Best regards,<br>
Min<br>
<div class=3D"x_x_x_x_x_moz-cite-prefix">On 2018/06/20 12: 14, Abu Naser wro=
te: <br>
</div>
<blockquote type=3D"cite">
<div id=3D"x_x_x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Jeff,</p>
<p><br>
</p>
<p>Yes, I am using a switch and other machines are also connected with=
that switch.
<br>
</p>
<p>If I remove other machines and just use my two node with the switch, the=
n will it improve the performance by 200 ~ 400 iterations?</p>
<p>Meanwhile I will give a try with a single dedicated cable. <span></span>=
<br>
</p>
<p><br>
</p>
<p>Thank you.<br>
</p>
<div id=3D"x_x_x_x_x_Signature">
<div id=3D"x_x_x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_x_x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Jeff Hammond <=
a class=3D"x_x_x_x_x_moz-txt-link-rfc2396E x_x_x_x_OWAAutoLink" href=3D"mai=
lto: jeff.science(a)gmail.com" id=3D"LPlnk983157">
<jeff.science(a)gmail.com></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 12:52:06 PM<br>
<b>To: </b> MPICH<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3Dutf-8">
<div>
<div dir=3D"ltr">Is the ethernet connection a single dedicated cable betwee=
n the two machines or are you running through a switch that handles other t=
raffic?
<div><br>
</div>
<div>My best guess is that this is noise and that you may be able to avoid =
it by running a very long time, e.g. 10000 iterations.</div>
<div><br>
</div>
<div>Jeff</div>
</div>
<div class=3D"x_x_x_x_x_x_gmail_extra"><br>
<div class=3D"x_x_x_x_x_x_gmail_quote">On Wed, Jun 20, 2018 at 6: 53 AM, Abu=
Naser <span dir=3D"ltr">
<<a href=3D"mailto: an16e(a)my.fsu.edu" target=3D"_blank" id=3D"LPlnk305789=
" class=3D"x_x_x_x_OWAAutoLink">an16e(a)my.fsu.edu</a>></span> wrote: <br>
<blockquote class=3D"x_x_x_x_x_x_gmail_quote">
<div dir=3D"ltr">
<div id=3D"x_x_x_x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"lt=
r">
<p><br>
</p>
<p>Good day to all,</p>
<p><br>
</p>
<p>I had run point to point osu_latency test in two nodes for 200 times.&nb=
sp; Followings are the average time in microsecond for various size of the =
messages -</p>
<div>1KB 84.8514 us<br>
<span>2KB 73.52535</span> us<br>
4KB 272.55275 us<br>
<span>8KB 234.86385</span> us<br>
16KB 288.88 us<br>
32KB 523.3725 us<br>
64KB 910.4025 us</div>
<p><br>
</p>
<p>From the above looks like, 2KB message has less latency than 1 KB and 8K=
B has less latency than 4KB.
<br>
</p>
<p>I was looking for explanation of this behavior but did not get any=
.</p>
<p><br>
</p>
<ol>
<li><span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span><span> is set to 128K=
B. So none of the above message size is using Rendezvous protocol. Is there=
any partition inside eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB -=
64KB)? If yes then what are the
boundaries for them? Can I log them with debug-event-logging? </span><br>
</li></ol>
<p><br>
</p>
<p>Setup I am using: </p>
<p>- two nodes has intel core i7, one with 16gb memory another one 8gb</p>
<p>- mpich 3.2.1, configured and build to use nemesis tcp</p>
<p>- 1gb Ethernet connection</p>
<p>- NFS is using for sharing<br>
</p>
<p>- osu_latency: uses MPI_Send and MPI_Recv</p>
<p>- <span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span>=3D <span>131072</sp=
an> (128KB)<br>
</p>
<p><br>
</p>
<p>Can anyone help me on that? Thanks in advance.<br>
</p>
<p><br>
</p>
<p><br>
</p>
<div id=3D"x_x_x_x_x_x_m_6077755676379859201Signature">
<div id=3D"x_x_x_x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"lt=
r">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
</div>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" id=3D"LPlnk816471" class=3D"x_x_x_x_OWAAutoLink">discuss(a)mpich.org</a><br=
>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank" id=3D"LPlnk624595" class=3D"x_x_x_x_OWAAutoLink">htt=
ps: //lists.mpich.org/<wbr>mailman/listinfo/discuss</a><br>
<br>
</blockquote>
</div>
<br>
<br>
<div><br>
</div>
-- <br>
<div class=3D"x_x_x_x_x_x_gmail_signature">Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank" id=3D"LPlnk3149=
93" class=3D"x_x_x_x_OWAAutoLink">jeff.science(a)gmail.com</a><br>
<a href=3D"http: //jeffhammond.github.io/" target=3D"_blank" id=3D"LPlnk8614=
34" class=3D"x_x_x_x_OWAAutoLink">http: //jeffhammond.github.io/</a></div>
</div>
</div>
<br>
<fieldset class=3D"x_x_x_x_x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________
discuss mailing list <a class=3D"x_x_x_x_x_moz-txt-link-abbreviated x_x=
_x_x_OWAAutoLink" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk657371">disc=
uss(a)mpich.org</a>
To manage subscription options or unsubscribe:
<a class=3D"x_x_x_x_x_moz-txt-link-freetext x_x_x_x_OWAAutoLink" href=3D"ht=
tps: //lists.mpich.org/mailman/listinfo/discuss" id=3D"LPlnk669988">https://=
lists.mpich.org/mailman/listinfo/discuss</a>
</pre>
</blockquote>
<br>
</div>
</div>
</div>
</div>
<br>
<fieldset class=3D"x_x_x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________
discuss mailing list <a class=3D"x_x_x_moz-txt-link-abbreviated x_x_OWA=
AutoLink" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk832953">discuss@mpic=
h.org</a>
To manage subscription options or unsubscribe:
<a class=3D"x_x_x_moz-txt-link-freetext x_x_OWAAutoLink" href=3D"https: //li=
sts.mpich.org/mailman/listinfo/discuss" id=3D"LPlnk481779">https: //lists.mp=
ich.org/mailman/listinfo/discuss</a>
</pre>
</blockquote>
<br>
</div>
</div>
<br>
<fieldset class=3D"x_x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________
discuss mailing list <a class=3D"x_x_moz-txt-link-abbreviated x_OWAAuto=
Link" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk408695">discuss(a)mpich.or=
g</a>
To manage subscription options or unsubscribe:
<a class=3D"x_x_moz-txt-link-freetext x_OWAAutoLink" href=3D"https: //lists.=
mpich.org/mailman/listinfo/discuss" id=3D"LPlnk572504">https: //lists.mpich.=
org/mailman/listinfo/discuss</a>
</pre>
</blockquote>
<br>
</div>
</div>
<br>
<fieldset class=3D"x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________
discuss mailing list <a class=3D"x_moz-txt-link-abbreviated" href=3D"ma=
ilto: discuss(a)mpich.org">discuss(a)mpich.org</a>
To manage subscription options or unsubscribe:
<a class=3D"x_moz-txt-link-freetext" href=3D"https: //lists.mpich.org/mailma=
n/listinfo/discuss">https: //lists.mpich.org/mailman/listinfo/discuss</a>
</pre>
</blockquote>
<br>
</div>
</body>
</html>
--_000_BLUPR0501MB2003D13F739833B52E59DBB097430BLUPR0501MB2003_--
--===============7802683794760088852==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============7802683794760088852==--
Message-ID: <sanitized-2315(a)migration.local>
1
0
as less latency than 4KB.
I was looking for explanation of this behavior but did not get any.
1. MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE is set to 128KB. So none of the abov=
e message size is using Rendezvous protocol. Is there any partition inside =
eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB)? If yes then wh=
at are the boundaries for them? Can I log them with debug-event-logging?
Setup I am using:
- two nodes has intel core i7, one with 16gb memory another one 8gb
- mpich 3.2.1, configured and build to use nemesis tcp
- 1gb Ethernet connection
- NFS is using for sharing
- osu_latency: uses MPI_Send and MPI_Recv
- MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE=3D 131072 (128KB)
Can anyone help me on that? Thanks in advance.
Best Regards,
Abu Naser
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--
Jeff Hammond
jeff.science(a)gmail.com<mailto: jeff.science(a)gmail.com>
http: //jeffhammond.github.io/
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--_000_BLUPR0501MB2003DCD7FDB382061050A6B997430BLUPR0501MB2003_
Content-Type: text/html; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
<html><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3DWindows-1=
252">
<style type=3D"text/css" style=3D"display: none;"><!-- P {margin-top:0;margi=
n-bottom: 0;} --></style>
</head>
<body dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt; color: rgb(0, 0,=
0); font-family: Calibri, Helvetica, sans-serif, "EmojiFont", &q=
uot;Apple Color Emoji", "Segoe UI Emoji", NotoColorEmoji, &q=
uot;Segoe UI Symbol", "Android Emoji", EmojiSymbols;" dir=3D=
"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt; color: rgb(0, 0,=
0); font-family: Calibri, Helvetica, sans-serif, "EmojiFont", &q=
uot;Apple Color Emoji", "Segoe UI Emoji", NotoColorEmoji, &q=
uot;Segoe UI Symbol", "Android Emoji", EmojiSymbols;" dir=3D=
"ltr">
<p style=3D"margin-top: 0;margin-bottom:0">Hello Min,</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">I have downloaded it from <a href=
=3D"http: //www.mpich.org/static/downloads/3.2.1/mpich-3.2.1.tar.gz" class=
=3D"OWAAutoLink" id=3D"LPlnk943697" previewremoved=3D"true">
http: //www.mpich.org/static/downloads/3.2.1/mpich-3.2.1.tar.gz</a> but it d=
id not work. I have received almost same error. Except this time no process=
information from my remote machine.</p>
<p style=3D"margin-top: 0;margin-bottom:0"><b>Previously I have received thi=
s - </b>
<br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"></p>
<div style=3D""><i><span style=3D"font-size: 10pt; color: rgb(255, 0, 0);">=
Process 3 of 4 is on dhcp16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt; color: rgb(255, 0, 0);">=
Process 1 of 4 is on dhcp16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt; color: rgb(255, 0, 0);">=
Process 0 of 4 is on dhcp16198</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt; color: rgb(255, 0, 0);">=
Process 2 of 4 is on dhcp16198</span></i></div>
<p></p>
<p style=3D"margin-top: 0;margin-bottom:0"><b>With the new source code -</b>=
</p>
<p style=3D"margin-top: 0;margin-bottom:0"></p>
<div><i><span style=3D"font-size: 10pt; color: rgb(255, 0, 0);">Process 0 o=
f 4 is on dhcp16198</span></i></div>
<div><i><span style=3D"font-size: 10pt; color: rgb(255, 0, 0);">Process 2 o=
f 4 is on dhcp16198</span></i></div>
<p></p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"><b>Entire error message is:</b></=
p>
<p style=3D"margin-top: 0;margin-bottom:0"></p>
<div><i><span style=3D"font-size: 10pt;">Process 0 of 4 is on dhcp16198</sp=
an></i></div>
<div><i><span style=3D"font-size: 10pt;">Process 2 of 4 is on dhcp16198</sp=
an></i></div>
<div><i><span style=3D"font-size: 10pt;">Fatal error in PMPI_Bcast: Unknown=
error class, error stack:</span></i></div>
<div><i><span style=3D"font-size: 10pt;">PMPI_Bcast(1600)..................=
..........: MPI_Bcast(buf=3D0x7ffd1ee145f0, count=3D1, MPI_INT, root=3D0, M=
PI_COMM_WORLD) failed</span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast_impl(1452).............=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast(1476)..................=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast_intra(1249)............=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_SMP_Bcast(1081)..............=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast_binomial(285)..........=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIC_Send(303)....................=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIC_Wait(226)....................=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIDI_CH3i_Progress_wait(242).....=
..........: an error occurred while handling an event returned by MPIDU_Soc=
k_Wait()</span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIDI_CH3I_Progress_handle_sock_ev=
ent(698)..: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIDI_CH3_Sockconn_handle_connect_=
event(597): [ch3:sock] failed to connnect to remote process</span></i></div=
>
<div><i><span style=3D"font-size: 10pt;">MPIDU_Socki_handle_connect(808)...=
..........: connection failure (set=3D0,sock=3D1,errno=3D111:Connection ref=
used)</span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_SMP_Bcast(1088)..............=
..........: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast_binomial(310)..........=
..........: Failure during collective</span></i></div>
<div><i><span style=3D"font-size: 10pt;">Fatal error in PMPI_Bcast: Other M=
PI error, error stack: </span></i></div>
<div><i><span style=3D"font-size: 10pt;">PMPI_Bcast(1600)........: MPI_Bcas=
t(buf=3D0x7ffe2eeb90f0, count=3D1, MPI_INT, root=3D0, MPI_COMM_WORLD) faile=
d</span></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast_impl(1452)...: </s=
pan></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast(1476)........: </s=
pan></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast_intra(1249)..: </s=
pan></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_SMP_Bcast(1088)....: </s=
pan></i></div>
<div><i><span style=3D"font-size: 10pt;">MPIR_Bcast_binomial(310): Failure =
during collective</span></i></div>
<br>
<p></p>
<p style=3D"margin-top: 0;margin-bottom:0">Again if I configure the new sour=
ce with tcp, it works fine.</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">Thank You.<br>
</p>
<div id=3D"Signature">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr style=3D"display: inline-block;width:98%" tabindex=3D"-1">
<div id=3D"divRplyFwdMsg" dir=3D"ltr"><font style=3D"font-size: 11pt" face=
=3D"Calibri, sans-serif" color=3D"#000000"><b>From: </b> Min Si <msi(a)anl.=
gov><br>
<b>Sent: </b> Monday, July 2, 2018 11:56:51 AM<br>
<b>To: </b> discuss(a)mpich.org<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Hi Abu,<br>
<br>
Thanks for reporting this. Can you please try the latest release with ch3/s=
ock and see if you still have this error ?
<br>
<br>
Min<br>
<div class=3D"x_moz-cite-prefix">On 2018/07/01 21: 47, Abu Naser wrote:<br>
</div>
<blockquote type=3D"cite">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; co=
lor: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-serif, "Emoji=
Font", "Apple Color Emoji", "Segoe UI Emoji", Noto=
ColorEmoji, "Segoe UI Symbol", "Android Emoji", EmojiSy=
mbols;">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p style=3D"">Hello Min,</p>
<p style=3D""><br>
</p>
<p style=3D"">After compiling my mpich-3.2.1 with sock, while I was trying =
to run any program including osu benchmark or examples/cpi&=
nbsp; in two machines, I have received following error -</p>
<p style=3D""><br>
</p>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 3 of 4 is on dhcp=
16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 1 of 4 is on dhcp=
16194</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 0 of 4 is on dhcp=
16198</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Process 2 of 4 is on dhcp=
16198</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Fatal error in PMPI_Bcast=
: Unknown error class, error stack:</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">PMPI_Bcast(1600).........=
...................: MPI_Bcast(buf=3D0x7ffc1808542c, count=3D1, MPI_INT, ro=
ot=3D0, MPI_COMM_WORLD) failed</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_impl(1452)....=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast(1476).........=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_intra(1249)...=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1081).....=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(285).=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIC_Send(303)...........=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIC_Wait(226)...........=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDI_CH3i_Progress_wait(=
242)...............: an error occurred while handling an event returned by =
MPIDU_Sock_Wait()</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDI_CH3I_Progress_handl=
e_sock_event(698)..: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDI_CH3_Sockconn_handle=
_connect_event(597): [ch3:sock] failed to connnect to remote process</span>=
</i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIDU_Socki_handle_connec=
t(808).............: connection failure (set=3D0,sock=3D1,errno=3D111:Conne=
ction refused)</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1088).....=
...................: </span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(310).=
...................: Failure during collective</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">Fatal error in PMPI_Bcast=
: Other MPI error, error stack:</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">PMPI_Bcast(1600)........:=
MPI_Bcast(buf=3D0x7ffd9eeebdac, count=3D1, MPI_INT, root=3D0, MPI_COMM_WOR=
LD) failed</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_impl(1452)...:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast(1476)........:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_intra(1249)..:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_SMP_Bcast(1088)....:=
</span></i></div>
<div style=3D""><i><span style=3D"font-size: 10pt">MPIR_Bcast_binomial(310):=
Failure during collective</span></i></div>
<br style=3D"">
<p style=3D""></p>
<p style=3D""><span style=3D"font-size: 12pt">I checked the mpich FAQ a=
nd also mpich discussion list. Based on that I have checked </span>fol=
lowings<span style=3D"font-size: 12pt"> </span><span style=3D"font-size=
: 12pt">and found they are fine in my machines -</span><br>
</p>
<p style=3D""><span style=3D"font-size: 12pt">- firewall is disabled in both=
machine</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- I can do </span>passwor=
d less<span style=3D"font-size: 12pt"> ssh in both machine</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- /etc/hosts in both machine c=
onfigured with ip address and name properly</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- I have updated the library p=
ath and used absolute path for mpiexec</span></p>
<p style=3D""><span style=3D"font-size: 12pt">- Most importantly when I conf=
igured and build mpich with tcp, it works fine.</span></p>
<p style=3D""><span style=3D"font-size: 12pt"><br>
</span></p>
<p style=3D""><span style=3D"font-size: 12pt"> I think I am </span=
><span style=3D"font-size: 12pt">missing something but could not figure out =
yet. Any help would be
</span>appreciated<span style=3D"font-size: 12pt">.</span></p>
<p style=3D""><span style=3D"font-size: 12pt"><br>
</span></p>
<p style=3D""><span style=3D"font-size: 12pt">Thank you.</span></p>
<br>
<p style=3D""><br>
</p>
<p style=3D""><br>
</p>
<p style=3D""><br>
</p>
<div id=3D"x_Signature" style=3D"">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr style=3D"display: inline-block; width:98%" tabindex=3D"-1">
<div id=3D"x_divRplyFwdMsg" dir=3D"ltr"><font style=3D"font-size: 11pt" face=
=3D"Calibri, sans-serif" color=3D"#000000"><b>From: </b> Min Si
<a class=3D"x_moz-txt-link-rfc2396E OWAAutoLink" href=3D"mailto: msi(a)anl.gov=
" id=3D"LPlnk191149" previewremoved=3D"true">
<msi(a)anl.gov></a><br>
<b>Sent: </b> Tuesday, June 26, 2018 12:54:29 PM<br>
<b>To: </b> <a class=3D"x_moz-txt-link-abbreviated OWAAutoLink" href=3D"mail=
to: discuss(a)mpich.org" id=3D"LPlnk414203" previewremoved=3D"true">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Hi Abu,<br>
<br>
I think the results are stable enough. Perhaps you could also try the follo=
wing tests, and see if similar trend exists: <br>
- MPICH/socket (set `--with-device=3Dch3: sock` at configure)<br>
- A socket-based pingpong test without MPI. <br>
<br>
At this point, I could not think of any MPI-specific design for 2k/8k messa=
ges. My guess is that it is related to your network connection.
<br>
<br>
Min<br>
<br>
<div class=3D"x_x_moz-cite-prefix">On 2018/06/24 11: 09, Abu Naser wrote:<br=
>
</div>
<blockquote type=3D"cite">
<meta content=3D"text/html; charset=3DWindows-1252">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Min and Jeff,</p>
<p><br>
</p>
<p>Here is my experiment results. Default number of iterations in osu_=
latency for 0B =96 8KB is 10,000. With that setting I had run the osu_laten=
cy 100 times and found standard deviation 33 for 8KB message size.</p>
<p><br>
</p>
<p>So later I have set the iteration to 50,000 and 100,000 for 1KB =96 16KB=
message size. Then run osu_latency for 100 times for each setting and take=
the average and standard deviation.</p>
<p><br>
</p>
<table width=3D"665">
<colgroup><col width=3D"99"><col width=3D"112"><col width=3D"118"><col widt=
h=3D"154"><col width=3D"140"></colgroup>
<tbody>
<tr>
<td width=3D"99">
<p><b>Msg Size in Bytes</b></p>
</td>
<td width=3D"112">
<p><b>Avg time in us (50K iterations)</b></p>
</td>
<td width=3D"118">
<p><b>Avg time in us (100k iterations)</b></p>
</td>
<td width=3D"154">
<p><b>Standard deviation (50K iterations)</b></p>
</td>
<td width=3D"140">
<p><b>Standard deviation (100K iterations)</b></p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>1k</p>
</td>
<td width=3D"112">
<p>85.10</p>
</td>
<td width=3D"118">
<p>84.9</p>
</td>
<td width=3D"154">
<p>0.55</p>
</td>
<td width=3D"140">
<p>0.45</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>2k</p>
</td>
<td width=3D"112">
<p>75.79</p>
</td>
<td width=3D"118">
<p>74.63</p>
</td>
<td width=3D"154">
<p>5.09</p>
</td>
<td width=3D"140">
<p>4.44</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>4k</p>
</td>
<td width=3D"112">
<p>273.80</p>
</td>
<td width=3D"118">
<p>274.71</p>
</td>
<td width=3D"154">
<p>4.18</p>
</td>
<td width=3D"140">
<p>2.45</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>8k</p>
</td>
<td width=3D"112">
<p>258.56</p>
</td>
<td width=3D"118">
<p>249.83</p>
</td>
<td width=3D"154">
<p>21.14</p>
</td>
<td width=3D"140">
<p>28</p>
</td>
</tr>
<tr>
<td width=3D"99" height=3D"24">
<p>16k</p>
</td>
<td width=3D"112">
<p>281.31</p>
</td>
<td width=3D"118">
<p>281.02</p>
</td>
<td width=3D"154">
<p>3.22</p>
</td>
<td width=3D"140">
<p>4.10</p>
</td>
</tr>
</tbody>
</table>
<p><br>
</p>
<p><br>
</p>
<p>The standard deviation of 8K message is so high and that implies it actu=
ally not producing any consistent latency time. Looks like that's the =
reason for 8K is taking less time than 4K.</p>
<p><br>
</p>
<p>Meanwhile, 2K has standard deviation less than 5 but 1K message latency =
timing are more densely populated than 2K. So probably this is the explanat=
ion for 2K message less latency time.</p>
<p><br>
</p>
<p>Thank you for your suggestions.</p>
<br>
<p><br>
</p>
<div id=3D"x_x_Signature">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Abu Naser<br>
<b>Sent: </b> Wednesday, June 20, 2018 1:48:53 PM<br>
<b>To: </b> <a class=3D"x_x_moz-txt-link-abbreviated x_OWAAutoLink" href=3D"=
mailto: discuss(a)mpich.org" id=3D"LPlnk729146" previewremoved=3D"true">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3Diso-8859-1">
<div dir=3D"ltr">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Min,</p>
<p><br>
</p>
<p>Thanks for the clarification. I will do the experiment.<br>
</p>
<p><br>
</p>
<div id=3D"x_x_x_Signature">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Thanks.</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Min Si <a class=3D=
"x_x_moz-txt-link-rfc2396E x_OWAAutoLink" href=3D"mailto: msi(a)anl.gov" id=3D=
"LPlnk558260" previewremoved=3D"true">
<msi(a)anl.gov></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 1:39:30 PM<br>
<b>To: </b> <a class=3D"x_x_moz-txt-link-abbreviated x_OWAAutoLink" href=3D"=
mailto: discuss(a)mpich.org" id=3D"LPlnk472728" previewremoved=3D"true">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div>Hi Abu,<br>
<br>
I think Jeff means that you should run your experiment with more iterations=
in order to get a stable results.<br>
- Increase the iteration of for loop in each execution (I think osu benchma=
rk allows you to set it)<br>
- Run the experiments 10 or 100 times, and take the average and standard de=
viation.<br>
<br>
If you see a very small standard deviation (e.g., <=3D5%), then the tren=
d is stable and you might not see such gaps.<br>
<br>
Best regards,<br>
Min<br>
<div class=3D"x_x_x_x_moz-cite-prefix">On 2018/06/20 12: 14, Abu Naser wrote=
: <br>
</div>
<blockquote type=3D"cite">
<div id=3D"x_x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Jeff,</p>
<p><br>
</p>
<p>Yes, I am using a switch and other machines are also connected with=
that switch.
<br>
</p>
<p>If I remove other machines and just use my two node with the switch, the=
n will it improve the performance by 200 ~ 400 iterations?</p>
<p>Meanwhile I will give a try with a single dedicated cable. <span></span>=
<br>
</p>
<p><br>
</p>
<p>Thank you.<br>
</p>
<div id=3D"x_x_x_x_Signature">
<div id=3D"x_x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Jeff Hammond <a =
class=3D"x_x_x_x_moz-txt-link-rfc2396E x_x_x_OWAAutoLink" href=3D"mailto: je=
ff.science(a)gmail.com" id=3D"LPlnk983157" previewremoved=3D"true">
<jeff.science(a)gmail.com></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 12:52:06 PM<br>
<b>To: </b> MPICH<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3Dutf-8">
<div>
<div dir=3D"ltr">Is the ethernet connection a single dedicated cable betwee=
n the two machines or are you running through a switch that handles other t=
raffic?
<div><br>
</div>
<div>My best guess is that this is noise and that you may be able to avoid =
it by running a very long time, e.g. 10000 iterations.</div>
<div><br>
</div>
<div>Jeff</div>
</div>
<div class=3D"x_x_x_x_x_gmail_extra"><br>
<div class=3D"x_x_x_x_x_gmail_quote">On Wed, Jun 20, 2018 at 6: 53 AM, Abu N=
aser <span dir=3D"ltr">
<<a href=3D"mailto: an16e(a)my.fsu.edu" target=3D"_blank" id=3D"LPlnk305789=
" class=3D"x_x_x_OWAAutoLink" previewremoved=3D"true">an16e(a)my.fsu.edu</a>&=
gt;</span> wrote: <br>
<blockquote class=3D"x_x_x_x_x_gmail_quote">
<div dir=3D"ltr">
<div id=3D"x_x_x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr"=
>
<p><br>
</p>
<p>Good day to all,</p>
<p><br>
</p>
<p>I had run point to point osu_latency test in two nodes for 200 times.&nb=
sp; Followings are the average time in microsecond for various size of the =
messages -</p>
<div>1KB 84.8514 us<br>
<span>2KB 73.52535</span> us<br>
4KB 272.55275 us<br>
<span>8KB 234.86385</span> us<br>
16KB 288.88 us<br>
32KB 523.3725 us<br>
64KB 910.4025 us</div>
<p><br>
</p>
<p>From the above looks like, 2KB message has less latency than 1 KB and 8K=
B has less latency than 4KB.
<br>
</p>
<p>I was looking for explanation of this behavior but did not get any=
.</p>
<p><br>
</p>
<ol>
<li><span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span><span> is set to 128K=
B. So none of the above message size is using Rendezvous protocol. Is there=
any partition inside eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB -=
64KB)? If yes then what are the
boundaries for them? Can I log them with debug-event-logging? </span><br>
</li></ol>
<p><br>
</p>
<p>Setup I am using: </p>
<p>- two nodes has intel core i7, one with 16gb memory another one 8gb</p>
<p>- mpich 3.2.1, configured and build to use nemesis tcp</p>
<p>- 1gb Ethernet connection</p>
<p>- NFS is using for sharing<br>
</p>
<p>- osu_latency: uses MPI_Send and MPI_Recv</p>
<p>- <span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span>=3D <span>131072</sp=
an> (128KB)<br>
</p>
<p><br>
</p>
<p>Can anyone help me on that? Thanks in advance.<br>
</p>
<p><br>
</p>
<p><br>
</p>
<div id=3D"x_x_x_x_x_m_6077755676379859201Signature">
<div id=3D"x_x_x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr"=
>
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
</div>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" id=3D"LPlnk816471" class=3D"x_x_x_OWAAutoLink" previewremoved=3D"true">di=
scuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank" id=3D"LPlnk624595" class=3D"x_x_x_OWAAutoLink" previ=
ewremoved=3D"true">https: //lists.mpich.org/<wbr>mailman/listinfo/discuss</a=
><br>
<br>
</blockquote>
</div>
<br>
<br>
<div><br>
</div>
-- <br>
<div class=3D"x_x_x_x_x_gmail_signature">Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank" id=3D"LPlnk3149=
93" class=3D"x_x_x_OWAAutoLink" previewremoved=3D"true">jeff.science(a)gmail.=
com</a><br>
<a href=3D"http: //jeffhammond.github.io/" target=3D"_blank" id=3D"LPlnk8614=
34" class=3D"x_x_x_OWAAutoLink" previewremoved=3D"true">http: //jeffhammond.=
github.io/</a></div>
</div>
</div>
<br>
<fieldset class=3D"x_x_x_x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________=0A=
discuss mailing list <a class=3D"x_x_x_x_moz-txt-link-abbreviated x_x_x=
_OWAAutoLink" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk657371" previewr=
emoved=3D"true">discuss(a)mpich.org</a>=0A=
To manage subscription options or unsubscribe: =0A=
<a class=3D"x_x_x_x_moz-txt-link-freetext x_x_x_OWAAutoLink" href=3D"https: =
//lists.mpich.org/mailman/listinfo/discuss" id=3D"LPlnk669988" previewremov=
ed=3D"true">https: //lists.mpich.org/mailman/listinfo/discuss</a>=0A=
</pre>
</blockquote>
<br>
</div>
</div>
</div>
</div>
<br>
<fieldset class=3D"x_x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________=0A=
discuss mailing list <a class=3D"x_x_moz-txt-link-abbreviated x_OWAAuto=
Link" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk832953" previewremoved=
=3D"true">discuss(a)mpich.org</a>=0A=
To manage subscription options or unsubscribe: =0A=
<a class=3D"x_x_moz-txt-link-freetext x_OWAAutoLink" href=3D"https: //lists.=
mpich.org/mailman/listinfo/discuss" id=3D"LPlnk481779" previewremoved=3D"tr=
ue">https: //lists.mpich.org/mailman/listinfo/discuss</a>=0A=
</pre>
</blockquote>
<br>
</div>
</div>
<br>
<fieldset class=3D"x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________=0A=
discuss mailing list <a class=3D"x_moz-txt-link-abbreviated OWAAutoLink=
" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk408695" previewremoved=3D"tr=
ue">discuss(a)mpich.org</a>=0A=
To manage subscription options or unsubscribe: =0A=
<a class=3D"x_moz-txt-link-freetext OWAAutoLink" href=3D"https: //lists.mpic=
h.org/mailman/listinfo/discuss" id=3D"LPlnk572504" previewremoved=3D"true">=
https: //lists.mpich.org/mailman/listinfo/discuss</a>=0A=
</pre>
</blockquote>
<br>
</div>
</div>
</body>
</html>
--_000_BLUPR0501MB2003DCD7FDB382061050A6B997430BLUPR0501MB2003_--
--===============0610270028680441889==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============0610270028680441889==--
Message-ID: <sanitized-2309(a)migration.local>
1
0
as less latency than 4KB.
I was looking for explanation of this behavior but did not get any.
1. MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE is set to 128KB. So none of the abov=
e message size is using Rendezvous protocol. Is there any partition inside =
eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB)? If yes then wh=
at are the boundaries for them? Can I log them with debug-event-logging?
Setup I am using:
- two nodes has intel core i7, one with 16gb memory another one 8gb
- mpich 3.2.1, configured and build to use nemesis tcp
- 1gb Ethernet connection
- NFS is using for sharing
- osu_latency: uses MPI_Send and MPI_Recv
- MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE=3D 131072 (128KB)
Can anyone help me on that? Thanks in advance.
Best Regards,
Abu Naser
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--
Jeff Hammond
jeff.science(a)gmail.com<mailto: jeff.science(a)gmail.com>
http: //jeffhammond.github.io/
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--_000_BLUPR0501MB2003414CB97CA97A0242D0BC97430BLUPR0501MB2003_
Content-Type: text/html; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
<html><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3DWindows-1=
252">
<style type=3D"text/css" style=3D"display: none;"><!-- P {margin-top:0;margi=
n-bottom: 0;} --></style>
</head>
<body dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt;color:#000000;font=
-family: Calibri,Helvetica,sans-serif;" dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"" dir=3D"ltr">
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt; margin-top: 0px; margin-bottom: 0px;">
Hello Min,</p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt; margin-top: 0px; margin-bottom: 0px;">
<br>
</p>
<p style=3D"margin-top: 0px; margin-bottom: 0px;"></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
After compiling my mpich-3.2.1 with sock, while I was trying to run a=
ny program including osu benchmark or examples/cpi in two m=
achines, I have received following error -</p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<br>
</p>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">Process 3 of 4 is on dhcp16194</span></=
i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">Process 1 of 4 is on dhcp16194</span></=
i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">Process 0 of 4 is on dhcp16198</span></=
i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">Process 2 of 4 is on dhcp16198</span></=
i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">Fatal error in PMPI_Bcast: Unknown erro=
r class, error stack: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">PMPI_Bcast(1600).......................=
.....: MPI_Bcast(buf=3D0x7ffc1808542c, count=3D1, MPI_INT, root=3D0, MPI_CO=
MM_WORLD) failed</span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast_impl(1452)..................=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast(1476).......................=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast_intra(1249).................=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_SMP_Bcast(1081)...................=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast_binomial(285)...............=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIC_Send(303).........................=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIC_Wait(226).........................=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIDI_CH3i_Progress_wait(242)..........=
.....: an error occurred while handling an event returned by MPIDU_Sock_Wai=
t()</span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIDI_CH3I_Progress_handle_sock_event(6=
98)..: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIDI_CH3_Sockconn_handle_connect_event=
(597): [ch3:sock] failed to connnect to remote process</span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIDU_Socki_handle_connect(808)........=
.....: connection failure (set=3D0,sock=3D1,errno=3D111:Connection refused)=
</span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_SMP_Bcast(1088)...................=
.....: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast_binomial(310)...............=
.....: Failure during collective</span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">Fatal error in PMPI_Bcast: Other MPI er=
ror, error stack: </span></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">PMPI_Bcast(1600)........: MPI_Bcast(buf=
=3D0x7ffd9eeebdac, count=3D1, MPI_INT, root=3D0, MPI_COMM_WORLD) failed</sp=
an></i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast_impl(1452)...: </span><=
/i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast(1476)........: </span><=
/i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast_intra(1249)..: </span><=
/i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_SMP_Bcast(1088)....: </span><=
/i></div>
<div style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-se=
rif, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", =
NotoColorEmoji, "Segoe UI Symbol", "Android Emoji", Emo=
jiSymbols; font-size: 16px;">
<i><span style=3D"font-size: 10pt;">MPIR_Bcast_binomial(310): Failure durin=
g collective</span></i></div>
<br style=3D"font-family: Calibri, Helvetica, sans-serif, EmojiFont, "=
Apple Color Emoji", "Segoe UI Emoji", NotoColorEmoji, "=
Segoe UI Symbol", "Android Emoji", EmojiSymbols; font-size: =
16px;">
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
</p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 16px;">
<span style=3D"font-size: 12pt;">I checked the mpich FAQ and also mpic=
h discussion list. Based on that I have checked </span>followings<span=
style=3D"font-size: 12pt;"> </span><span style=3D"font-size: 12pt;">a=
nd found they are fine in my machines -</span><br>
</p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<span style=3D"font-size: 12pt;">- firewall is disabled in both machine</sp=
an></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<span style=3D"font-size: 12pt;">- I can do </span>password less<span =
style=3D"font-size: 12pt;"> ssh in both machine</span></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<span style=3D"font-size: 12pt;">- /etc/hosts in both machine configured wi=
th ip address and name properly</span></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<span style=3D"font-size: 12pt;">- I have updated the library path and used=
absolute path for mpiexec</span></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<span style=3D"font-size: 12pt;">- Most importantly when I configured and b=
uild mpich with tcp, it works fine.</span></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<span style=3D"font-size: 12pt;"><br>
</span></p>
<p style=3D""><span style=3D"font-size: 12pt;"> I think I am </sp=
an><span style=3D"font-size: 12pt;">missing something but could not figure =
out yet. Any help would be
</span>appreciated<span style=3D"font-size: 12pt;">.</span></p>
<p style=3D""><span style=3D"font-size: 12pt;"><br>
</span></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt;">
<span style=3D"font-size: 12pt;">Thank you.</span></p>
<br>
<p></p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt; margin-top: 0px; margin-bottom: 0px;">
<br>
</p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt; margin-top: 0px; margin-bottom: 0px;">
<br>
</p>
<p style=3D"color: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-seri=
f, EmojiFont, "Apple Color Emoji", "Segoe UI Emoji", No=
toColorEmoji, "Segoe UI Symbol", "Android Emoji", Emoji=
Symbols; font-size: 12pt; margin-top: 0px; margin-bottom: 0px;">
<br>
</p>
<div id=3D"Signature" style=3D"color: rgb(0, 0, 0); font-family: Calibri, H=
elvetica, sans-serif, EmojiFont, "Apple Color Emoji", "Segoe=
UI Emoji", NotoColorEmoji, "Segoe UI Symbol", "Android=
Emoji", EmojiSymbols; font-size: 12pt;">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr style=3D"display: inline-block;width:98%" tabindex=3D"-1">
<div id=3D"divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif" st=
yle=3D"font-size: 11pt" color=3D"#000000"><b>From:</b> Min Si <msi(a)anl.go=
v><br>
<b>Sent: </b> Tuesday, June 26, 2018 12:54:29 PM<br>
<b>To: </b> discuss(a)mpich.org<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Hi Abu,<br>
<br>
I think the results are stable enough. Perhaps you could also try the follo=
wing tests, and see if similar trend exists: <br>
- MPICH/socket (set `--with-device=3Dch3: sock` at configure)<br>
- A socket-based pingpong test without MPI. <br>
<br>
At this point, I could not think of any MPI-specific design for 2k/8k messa=
ges. My guess is that it is related to your network connection.
<br>
<br>
Min<br>
<br>
<div class=3D"x_moz-cite-prefix">On 2018/06/24 11: 09, Abu Naser wrote:<br>
</div>
<blockquote type=3D"cite">
<meta content=3D"text/html;=0A=
charset=3DWindows-1252">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Min and Jeff,</p>
<p><br>
</p>
<p>Here is my experiment results. Default number of iterations in osu_=
latency for 0B =96 8KB is 10,000. With that setting I had run the osu_laten=
cy 100 times and found standard deviation 33 for 8KB message size.</p>
<p><br>
</p>
<p>So later I have set the iteration to 50,000 and 100,000 for 1KB =96 16KB=
message size. Then run osu_latency for 100 times for each setting and take=
the average and standard deviation.</p>
<p><br>
</p>
<table width=3D"665">
<colgroup><col width=3D"99"><col width=3D"112"><col width=3D"118"><col widt=
h=3D"154"><col width=3D"140"></colgroup>
<tbody>
<tr>
<td width=3D"99">
<p><b>Msg Size in Bytes</b></p>
</td>
<td width=3D"112">
<p><b>Avg time in us (50K iterations)</b></p>
</td>
<td width=3D"118">
<p><b>Avg time in us (100k iterations)</b></p>
</td>
<td width=3D"154">
<p><b>Standard deviation (50K iterations)</b></p>
</td>
<td width=3D"140">
<p><b>Standard deviation (100K iterations)</b></p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>1k</p>
</td>
<td width=3D"112">
<p>85.10</p>
</td>
<td width=3D"118">
<p>84.9</p>
</td>
<td width=3D"154">
<p>0.55</p>
</td>
<td width=3D"140">
<p>0.45</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>2k</p>
</td>
<td width=3D"112">
<p>75.79</p>
</td>
<td width=3D"118">
<p>74.63</p>
</td>
<td width=3D"154">
<p>5.09</p>
</td>
<td width=3D"140">
<p>4.44</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>4k</p>
</td>
<td width=3D"112">
<p>273.80</p>
</td>
<td width=3D"118">
<p>274.71</p>
</td>
<td width=3D"154">
<p>4.18</p>
</td>
<td width=3D"140">
<p>2.45</p>
</td>
</tr>
<tr>
<td width=3D"99">
<p>8k</p>
</td>
<td width=3D"112">
<p>258.56</p>
</td>
<td width=3D"118">
<p>249.83</p>
</td>
<td width=3D"154">
<p>21.14</p>
</td>
<td width=3D"140">
<p>28</p>
</td>
</tr>
<tr>
<td height=3D"24" width=3D"99">
<p>16k</p>
</td>
<td width=3D"112">
<p>281.31</p>
</td>
<td width=3D"118">
<p>281.02</p>
</td>
<td width=3D"154">
<p>3.22</p>
</td>
<td width=3D"140">
<p>4.10</p>
</td>
</tr>
</tbody>
</table>
<p><br>
</p>
<p><br>
</p>
<p>The standard deviation of 8K message is so high and that implies it actu=
ally not producing any consistent latency time. Looks like that's the =
reason for 8K is taking less time than 4K.</p>
<p><br>
</p>
<p>Meanwhile, 2K has standard deviation less than 5 but 1K message latency =
timing are more densely populated than 2K. So probably this is the explanat=
ion for 2K message less latency time.</p>
<p><br>
</p>
<p>Thank you for your suggestions.</p>
<br>
<p><br>
</p>
<div id=3D"x_Signature">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Abu Naser<br>
<b>Sent: </b> Wednesday, June 20, 2018 1:48:53 PM<br>
<b>To: </b> <a class=3D"x_moz-txt-link-abbreviated OWAAutoLink" href=3D"mail=
to: discuss(a)mpich.org" id=3D"LPlnk729146" previewremoved=3D"true">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3Diso-8859-1">
<div dir=3D"ltr">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Min,</p>
<p><br>
</p>
<p>Thanks for the clarification. I will do the experiment.<br>
</p>
<p><br>
</p>
<div id=3D"x_x_Signature">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Thanks.</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Min Si <a class=3D"x=
_moz-txt-link-rfc2396E OWAAutoLink" href=3D"mailto: msi(a)anl.gov" id=3D"LPlnk=
558260" previewremoved=3D"true">
<msi(a)anl.gov></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 1:39:30 PM<br>
<b>To: </b> <a class=3D"x_moz-txt-link-abbreviated OWAAutoLink" href=3D"mail=
to: discuss(a)mpich.org" id=3D"LPlnk472728" previewremoved=3D"true">
discuss(a)mpich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div>Hi Abu,<br>
<br>
I think Jeff means that you should run your experiment with more iterations=
in order to get a stable results.<br>
- Increase the iteration of for loop in each execution (I think osu benchma=
rk allows you to set it)<br>
- Run the experiments 10 or 100 times, and take the average and standard de=
viation.<br>
<br>
If you see a very small standard deviation (e.g., <=3D5%), then the tren=
d is stable and you might not see such gaps.<br>
<br>
Best regards,<br>
Min<br>
<div class=3D"x_x_x_moz-cite-prefix">On 2018/06/20 12: 14, Abu Naser wrote:<=
br>
</div>
<blockquote type=3D"cite">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p>Hello Jeff,</p>
<p><br>
</p>
<p>Yes, I am using a switch and other machines are also connected with=
that switch.
<br>
</p>
<p>If I remove other machines and just use my two node with the switch, the=
n will it improve the performance by 200 ~ 400 iterations?</p>
<p>Meanwhile I will give a try with a single dedicated cable. <span></span>=
<br>
</p>
<p><br>
</p>
<p>Thank you.<br>
</p>
<div id=3D"x_x_x_Signature">
<div id=3D"x_x_x_divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1">
<div id=3D"x_x_x_divRplyFwdMsg" dir=3D"ltr"><b>From: </b> Jeff Hammond <a cl=
ass=3D"x_x_x_moz-txt-link-rfc2396E x_x_OWAAutoLink" href=3D"mailto: jeff.sci=
ence(a)gmail.com" id=3D"LPlnk983157" previewremoved=3D"true">
<jeff.science(a)gmail.com></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 12:52:06 PM<br>
<b>To: </b> MPICH<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?
<div> </div>
</div>
<meta content=3D"text/html; charset=3Dutf-8">
<div>
<div dir=3D"ltr">Is the ethernet connection a single dedicated cable betwee=
n the two machines or are you running through a switch that handles other t=
raffic?
<div><br>
</div>
<div>My best guess is that this is noise and that you may be able to avoid =
it by running a very long time, e.g. 10000 iterations.</div>
<div><br>
</div>
<div>Jeff</div>
</div>
<div class=3D"x_x_x_x_gmail_extra"><br>
<div class=3D"x_x_x_x_gmail_quote">On Wed, Jun 20, 2018 at 6: 53 AM, Abu Nas=
er <span dir=3D"ltr">
<<a href=3D"mailto: an16e(a)my.fsu.edu" target=3D"_blank" id=3D"LPlnk305789=
" class=3D"x_x_OWAAutoLink" previewremoved=3D"true">an16e(a)my.fsu.edu</a>>=
;</span> wrote: <br>
<blockquote class=3D"x_x_x_x_gmail_quote">
<div dir=3D"ltr">
<div id=3D"x_x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p>Good day to all,</p>
<p><br>
</p>
<p>I had run point to point osu_latency test in two nodes for 200 times.&nb=
sp; Followings are the average time in microsecond for various size of the =
messages -</p>
<div>1KB 84.8514 us<br>
<span>2KB 73.52535</span> us<br>
4KB 272.55275 us<br>
<span>8KB 234.86385</span> us<br>
16KB 288.88 us<br>
32KB 523.3725 us<br>
64KB 910.4025 us</div>
<p><br>
</p>
<p>From the above looks like, 2KB message has less latency than 1 KB and 8K=
B has less latency than 4KB.
<br>
</p>
<p>I was looking for explanation of this behavior but did not get any=
.</p>
<p><br>
</p>
<ol>
<li><span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span><span> is set to 128K=
B. So none of the above message size is using Rendezvous protocol. Is there=
any partition inside eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB -=
64KB)? If yes then what are the
boundaries for them? Can I log them with debug-event-logging? </span><br>
</li></ol>
<p><br>
</p>
<p>Setup I am using: </p>
<p>- two nodes has intel core i7, one with 16gb memory another one 8gb</p>
<p>- mpich 3.2.1, configured and build to use nemesis tcp</p>
<p>- 1gb Ethernet connection</p>
<p>- NFS is using for sharing<br>
</p>
<p>- osu_latency: uses MPI_Send and MPI_Recv</p>
<p>- <span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span>=3D <span>131072</sp=
an> (128KB)<br>
</p>
<p><br>
</p>
<p>Can anyone help me on that? Thanks in advance.<br>
</p>
<p><br>
</p>
<p><br>
</p>
<div id=3D"x_x_x_x_m_6077755676379859201Signature">
<div id=3D"x_x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr">
<p><br>
</p>
<p><span>Best Regards,</span></p>
<span></span>
<div><span></span></div>
<span></span>
<p><span>Abu Naser</span><br>
</p>
</div>
</div>
</div>
</div>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" id=3D"LPlnk816471" class=3D"x_x_OWAAutoLink" previewremoved=3D"true">disc=
uss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank" id=3D"LPlnk624595" class=3D"x_x_OWAAutoLink" preview=
removed=3D"true">https: //lists.mpich.org/<wbr>mailman/listinfo/discuss</a><=
br>
<br>
</blockquote>
</div>
<br>
<br>
<div><br>
</div>
-- <br>
<div class=3D"x_x_x_x_gmail_signature">Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank" id=3D"LPlnk3149=
93" class=3D"x_x_OWAAutoLink" previewremoved=3D"true">jeff.science(a)gmail.co=
m</a><br>
<a href=3D"http: //jeffhammond.github.io/" target=3D"_blank" id=3D"LPlnk8614=
34" class=3D"x_x_OWAAutoLink" previewremoved=3D"true">http: //jeffhammond.gi=
thub.io/</a></div>
</div>
</div>
<br>
<fieldset class=3D"x_x_x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________=0A=
discuss mailing list <a class=3D"x_x_x_moz-txt-link-abbreviated x_x_OWA=
AutoLink" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk657371" previewremov=
ed=3D"true">discuss(a)mpich.org</a>=0A=
To manage subscription options or unsubscribe: =0A=
<a class=3D"x_x_x_moz-txt-link-freetext x_x_OWAAutoLink" href=3D"https: //li=
sts.mpich.org/mailman/listinfo/discuss" id=3D"LPlnk669988" previewremoved=
=3D"true">https: //lists.mpich.org/mailman/listinfo/discuss</a>=0A=
</pre>
</blockquote>
<br>
</div>
</div>
</div>
</div>
<br>
<fieldset class=3D"x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________=0A=
discuss mailing list <a class=3D"x_moz-txt-link-abbreviated OWAAutoLink=
" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk832953" previewremoved=3D"tr=
ue">discuss(a)mpich.org</a>=0A=
To manage subscription options or unsubscribe: =0A=
<a class=3D"x_moz-txt-link-freetext OWAAutoLink" href=3D"https: //lists.mpic=
h.org/mailman/listinfo/discuss" id=3D"LPlnk481779" previewremoved=3D"true">=
https: //lists.mpich.org/mailman/listinfo/discuss</a>=0A=
</pre>
</blockquote>
<br>
</div>
</div>
</body>
</html>
--_000_BLUPR0501MB2003414CB97CA97A0242D0BC97430BLUPR0501MB2003_--
--===============7322407779089830927==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============7322407779089830927==--
Message-ID: <sanitized-2306(a)migration.local>
1
0
as less latency than 4KB.
I was looking for explanation of this behavior but did not get any.
1. MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE is set to 128KB. So none of the abov=
e message size is using Rendezvous protocol. Is there any partition inside =
eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB)? If yes then wh=
at are the boundaries for them? Can I log them with debug-event-logging?
Setup I am using:
- two nodes has intel core i7, one with 16gb memory another one 8gb
- mpich 3.2.1, configured and build to use nemesis tcp
- 1gb Ethernet connection
- NFS is using for sharing
- osu_latency: uses MPI_Send and MPI_Recv
- MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE=3D 131072 (128KB)
Can anyone help me on that? Thanks in advance.
Best Regards,
Abu Naser
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--
Jeff Hammond
jeff.science(a)gmail.com<mailto: jeff.science(a)gmail.com>
http: //jeffhammond.github.io/
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--_000_BLUPR0501MB20037286A10BAAD88A18EF8D974B0BLUPR0501MB2003_
Content-Type: text/html; charset="Windows-1252"
Content-Transfer-Encoding: quoted-printable
<html><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3DWindows-1=
252">
<style type=3D"text/css" style=3D"display: none;"><!-- P {margin-top:0;margi=
n-bottom: 0;} --></style>
</head>
<body dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt;color:#000000;font=
-family: Calibri,Helvetica,sans-serif;" dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt; color: rgb(0, 0,=
0); font-family: Calibri, Helvetica, sans-serif, EmojiFont, "Apple Co=
lor Emoji", "Segoe UI Emoji", NotoColorEmoji, "Segoe UI=
Symbol", "Android Emoji", EmojiSymbols;" dir=3D"ltr">
<p style=3D"margin-top: 0;margin-bottom:0">Hello Min and Jeff,</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"></p>
<p style=3D"margin-bottom: 0in; line-height: 100%">Here is my experiment re=
sults. Default number of iterations in osu_latency for 0B =96 8KB is 1=
0,000. With that setting I had run the osu_latency 100 times and found stan=
dard deviation 33 for 8KB message size.</p>
<p style=3D"margin-bottom: 0in; line-height: 100%"><br>
</p>
<p style=3D"margin-bottom: 0in; line-height: 100%">So later I have set the =
iteration to 50,000 and 100,000 for 1KB =96 16KB message size. Then run osu=
_latency for 100 times for each setting and take the average and standard d=
eviation.</p>
<p style=3D"margin-bottom: 0in; line-height: 100%"><br>
</p>
<table width=3D"665" cellpadding=3D"4" cellspacing=3D"0">
<colgroup><col width=3D"99"><col width=3D"112"><col width=3D"118"><col widt=
h=3D"154"><col width=3D"140"></colgroup>
<tbody>
<tr valign=3D"top">
<td width=3D"99" style=3D"border-top: 1px solid #000000; border-bottom: 1px=
solid #000000; border-left: 1px solid #000000; border-right: none; padding=
-top: 0.04in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right: =
0in">
<p align=3D"center"><b>Msg Size in Bytes</b></p>
</td>
<td width=3D"112" style=3D"border-top: 1px solid #000000; border-bottom: 1p=
x solid #000000; border-left: 1px solid #000000; border-right: none; paddin=
g-top: 0.04in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right:=
0in">
<p align=3D"center"><b>Avg time in us (50K iterations)</b></p>
</td>
<td width=3D"118" style=3D"border-top: 1px solid #000000; border-bottom: 1p=
x solid #000000; border-left: 1px solid #000000; border-right: none; paddin=
g-top: 0.04in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right:=
0in">
<p align=3D"center"><b>Avg time in us (100k iterations)</b></p>
</td>
<td width=3D"154" style=3D"border-top: 1px solid #000000; border-bottom: 1p=
x solid #000000; border-left: 1px solid #000000; border-right: none; paddin=
g-top: 0.04in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right:=
0in">
<p align=3D"center"><b>Standard deviation (50K iterations)</b></p>
</td>
<td width=3D"140" style=3D"border: 1px solid #000000; padding: 0.04in">
<p align=3D"center"><b>Standard deviation (100K iterations)</b></p>
</td>
</tr>
<tr valign=3D"top">
<td width=3D"99" style=3D"border-top: none; border-bottom: 1px solid #00000=
0; border-left: 1px solid #000000; border-right: none; padding-top: 0in; pa=
dding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">1k</p>
</td>
<td width=3D"112" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">85.10</p>
</td>
<td width=3D"118" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">84.9</p>
</td>
<td width=3D"154" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">0.55</p>
</td>
<td width=3D"140" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: 1px solid #000000; paddin=
g-top: 0in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0.=
04in">
<p align=3D"center">0.45</p>
</td>
</tr>
<tr valign=3D"top">
<td width=3D"99" style=3D"border-top: none; border-bottom: 1px solid #00000=
0; border-left: 1px solid #000000; border-right: none; padding-top: 0in; pa=
dding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">2k</p>
</td>
<td width=3D"112" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">75.79</p>
</td>
<td width=3D"118" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">74.63</p>
</td>
<td width=3D"154" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center"><font color=3D"#ff0000">5.09</font></p>
</td>
<td width=3D"140" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: 1px solid #000000; paddin=
g-top: 0in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0.=
04in">
<p align=3D"center"><font color=3D"#ff0000">4.44</font></p>
</td>
</tr>
<tr valign=3D"top">
<td width=3D"99" style=3D"border-top: none; border-bottom: 1px solid #00000=
0; border-left: 1px solid #000000; border-right: none; padding-top: 0in; pa=
dding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">4k</p>
</td>
<td width=3D"112" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">273.80</p>
</td>
<td width=3D"118" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">274.71</p>
</td>
<td width=3D"154" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">4.18</p>
</td>
<td width=3D"140" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: 1px solid #000000; paddin=
g-top: 0in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0.=
04in">
<p align=3D"center">2.45</p>
</td>
</tr>
<tr valign=3D"top">
<td width=3D"99" style=3D"border-top: none; border-bottom: 1px solid #00000=
0; border-left: 1px solid #000000; border-right: none; padding-top: 0in; pa=
dding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">8k</p>
</td>
<td width=3D"112" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">258.56</p>
</td>
<td width=3D"118" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">249.83</p>
</td>
<td width=3D"154" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center"><font color=3D"#ff0000">21.14</font></p>
</td>
<td width=3D"140" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: 1px solid #000000; paddin=
g-top: 0in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0.=
04in">
<p align=3D"center"><font color=3D"#ff0000">28</font></p>
</td>
</tr>
<tr valign=3D"top">
<td width=3D"99" height=3D"24" style=3D"border-top: none; border-bottom: 1p=
x solid #000000; border-left: 1px solid #000000; border-right: none; paddin=
g-top: 0in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0i=
n">
<p align=3D"center">16k</p>
</td>
<td width=3D"112" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">281.31</p>
</td>
<td width=3D"118" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">281.02</p>
</td>
<td width=3D"154" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: none; padding-top: 0in; p=
adding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0in">
<p align=3D"center">3.22</p>
</td>
<td width=3D"140" style=3D"border-top: none; border-bottom: 1px solid #0000=
00; border-left: 1px solid #000000; border-right: 1px solid #000000; paddin=
g-top: 0in; padding-bottom: 0.04in; padding-left: 0.04in; padding-right: 0.=
04in">
<p align=3D"center">4.10</p>
</td>
</tr>
</tbody>
</table>
<p style=3D"margin-bottom: 0in; line-height: 100%"><br>
</p>
<p style=3D"margin-bottom: 0in; line-height: 100%"><br>
</p>
<p style=3D"margin-bottom: 0in; line-height: 100%">The standard deviation o=
f 8K message is so high and that implies it actually not producing any cons=
istent latency time. Looks like that's the reason for 8K is taking les=
s time than 4K.</p>
<p style=3D"margin-bottom: 0in; line-height: 100%"><br>
</p>
<p style=3D"margin-bottom: 0in; line-height: 100%">Meanwhile, 2K has standa=
rd deviation less than 5 but 1K message latency timing are more densely pop=
ulated than 2K. So probably this is the explanation for 2K message less lat=
ency time.</p>
<p style=3D"margin-bottom: 0in; line-height: 100%"><br>
</p>
<p style=3D"margin-bottom: 0in; line-height: 100%">Thank you for your sugge=
stions.</p>
<br>
<p></p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<div id=3D"Signature">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr style=3D"display: inline-block;width:98%" tabindex=3D"-1">
<div id=3D"divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif" st=
yle=3D"font-size: 11pt" color=3D"#000000"><b>From:</b> Abu Naser<br>
<b>Sent: </b> Wednesday, June 20, 2018 1:48:53 PM<br>
<b>To: </b> discuss(a)mpich.org<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3Diso-8859-1">
<div dir=3D"ltr">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; co=
lor: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-serif, EmojiFont, =
"Apple Color Emoji", "Segoe UI Emoji", NotoColorEmoji, =
"Segoe UI Symbol", "Android Emoji", EmojiSymbols;">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; col=
or: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont&quo=
t;,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,=
"Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p style=3D"margin-top: 0; margin-bottom:0">Hello Min,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Thanks for the clarification.&nb=
sp; I will do the experiment.<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<div id=3D"x_Signature">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; col=
or: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont&quo=
t;,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,=
"Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p>Thanks.</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1" style=3D"display: inline-block; width:98%">
<div id=3D"x_divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif" =
color=3D"#000000" style=3D"font-size: 11pt"><b>From:</b> Min Si <msi(a)anl.=
gov><br>
<b>Sent: </b> Wednesday, June 20, 2018 1:39:30 PM<br>
<b>To: </b> discuss(a)mpich.org<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Hi Abu,<br>
<br>
I think Jeff means that you should run your experiment with more iterations=
in order to get a stable results.<br>
- Increase the iteration of for loop in each execution (I think osu benchma=
rk allows you to set it)<br>
- Run the experiments 10 or 100 times, and take the average and standard de=
viation.<br>
<br>
If you see a very small standard deviation (e.g., <=3D5%), then the tren=
d is stable and you might not see such gaps.<br>
<br>
Best regards,<br>
Min<br>
<div class=3D"x_x_moz-cite-prefix">On 2018/06/20 12: 14, Abu Naser wrote:<br=
>
</div>
<blockquote type=3D"cite">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; c=
olor: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont&q=
uot;,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoj=
i,"Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p style=3D"margin-top: 0; margin-bottom:0">Hello Jeff,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Yes, I am using a switch and oth=
er machines are also connected with that switch.
<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">If I remove other machines and j=
ust use my two node with the switch, then will it improve the performance b=
y 200 ~ 400 iterations?</p>
<p style=3D"margin-top: 0; margin-bottom:0">Meanwhile I will give a try with=
a single dedicated cable.
<span></span><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Thank you.<br>
</p>
<div id=3D"x_x_Signature">
<div id=3D"x_x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1" style=3D"display: inline-block; width:98%">
<div id=3D"x_x_divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif=
" color=3D"#000000" style=3D"font-size: 11pt"><b>From:</b> Jeff Hammond
<a class=3D"x_x_moz-txt-link-rfc2396E x_OWAAutoLink" href=3D"mailto: jeff.sc=
ience(a)gmail.com" id=3D"LPlnk983157" previewremoved=3D"true">
<jeff.science(a)gmail.com></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 12:52:06 PM<br>
<b>To: </b> MPICH<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3Dutf-8">
<div>
<div dir=3D"ltr">Is the ethernet connection a single dedicated cable betwee=
n the two machines or are you running through a switch that handles other t=
raffic?
<div><br>
</div>
<div>My best guess is that this is noise and that you may be able to avoid =
it by running a very long time, e.g. 10000 iterations.</div>
<div><br>
</div>
<div>Jeff</div>
</div>
<div class=3D"x_x_x_gmail_extra"><br>
<div class=3D"x_x_x_gmail_quote">On Wed, Jun 20, 2018 at 6: 53 AM, Abu Naser=
<span dir=3D"ltr">
<<a href=3D"mailto: an16e(a)my.fsu.edu" target=3D"_blank" id=3D"LPlnk305789=
" class=3D"x_OWAAutoLink" previewremoved=3D"true">an16e(a)my.fsu.edu</a>><=
/span> wrote: <br>
<blockquote class=3D"x_x_x_gmail_quote" style=3D"margin: 0 0 0 .8ex; border-=
left: 1px #ccc solid; padding-left:1ex">
<div dir=3D"ltr">
<div id=3D"x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr" sty=
le=3D"">
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Good day to all,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I had run point to point osu_lat=
ency test in two nodes for 200 times. Followings are the average time=
in microsecond for various size of the messages -</p>
<div>1KB 84.8514 us<br>
<span style=3D"color: rgb(255,0,0)">2KB 73.52535</span> us=
<br>
4KB 272.55275 us<br>
<span style=3D"color: rgb(255,0,0)">8KB 234.86385</span> u=
s<br>
16KB 288.88 us<br>
32KB 523.3725 us<br>
64KB 910.4025 us</div>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">From the above looks like, 2KB m=
essage has less latency than 1 KB and 8KB has less latency than 4KB.
<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I was looking for explanation of=
this behavior but did not get any.</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<ol style=3D"margin-bottom: 0px; margin-top:0px">
<li><span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span><span> is set to 128K=
B. So none of the above message size is using Rendezvous protocol. Is there=
any partition inside eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB -=
64KB)? If yes then what are the
boundaries for them? Can I log them with debug-event-logging? </span><br>
</li></ol>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Setup I am using:</p>
<p style=3D"margin-top: 0; margin-bottom:0">- two nodes has intel core i7, o=
ne with 16gb memory another one 8gb</p>
<p style=3D"margin-top: 0; margin-bottom:0">- mpich 3.2.1, configured and bu=
ild to use nemesis tcp</p>
<p style=3D"margin-top: 0; margin-bottom:0">- 1gb Ethernet connection</p>
<p style=3D"margin-top: 0; margin-bottom:0">- NFS is using for sharing<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">- osu_latency : uses MPI_Send an=
d MPI_Recv</p>
<p style=3D"margin-top: 0; margin-bottom:0">- <span>MPIR_CVAR_CH3_EAGER_MAX_=
MSG_<wbr>SIZE</span>=3D
<span>131072</span> (128KB)<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Can anyone help me on that? Than=
ks in advance.<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<div id=3D"x_x_x_m_6077755676379859201Signature">
<div id=3D"x_x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr" sty=
le=3D"">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
</div>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" id=3D"LPlnk816471" class=3D"x_OWAAutoLink" previewremoved=3D"true">discus=
s(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank" id=3D"LPlnk624595" class=3D"x_OWAAutoLink" previewre=
moved=3D"true">https: //lists.mpich.org/<wbr>mailman/listinfo/discuss</a><br=
>
<br>
</blockquote>
</div>
<br>
<br clear=3D"all">
<div><br>
</div>
-- <br>
<div class=3D"x_x_x_gmail_signature">Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank" id=3D"LPlnk3149=
93" class=3D"x_OWAAutoLink" previewremoved=3D"true">jeff.science(a)gmail.com<=
/a><br>
<a href=3D"http: //jeffhammond.github.io/" target=3D"_blank" id=3D"LPlnk8614=
34" class=3D"x_OWAAutoLink" previewremoved=3D"true">http: //jeffhammond.gith=
ub.io/</a></div>
</div>
</div>
<br>
<fieldset class=3D"x_x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________=0A=
discuss mailing list <a class=3D"x_x_moz-txt-link-abbreviated x_OWAAuto=
Link" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk657371" previewremoved=
=3D"true">discuss(a)mpich.org</a>=0A=
To manage subscription options or unsubscribe: =0A=
<a class=3D"x_x_moz-txt-link-freetext x_OWAAutoLink" href=3D"https: //lists.=
mpich.org/mailman/listinfo/discuss" id=3D"LPlnk669988" previewremoved=3D"tr=
ue">https: //lists.mpich.org/mailman/listinfo/discuss</a>=0A=
</pre>
</blockquote>
<br>
</div>
</div>
</div>
</div>
</body>
</html>
--_000_BLUPR0501MB20037286A10BAAD88A18EF8D974B0BLUPR0501MB2003_--
--===============3923392798546736373==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============3923392798546736373==--
Message-ID: <sanitized-2285(a)migration.local>
1
0
as less latency than 4KB.
I was looking for explanation of this behavior but did not get any.
1. MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE is set to 128KB. So none of the abov=
e message size is using Rendezvous protocol. Is there any partition inside =
eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB)? If yes then wh=
at are the boundaries for them? Can I log them with debug-event-logging?
Setup I am using:
- two nodes has intel core i7, one with 16gb memory another one 8gb
- mpich 3.2.1, configured and build to use nemesis tcp
- 1gb Ethernet connection
- NFS is using for sharing
- osu_latency: uses MPI_Send and MPI_Recv
- MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE=3D 131072 (128KB)
Can anyone help me on that? Thanks in advance.
Best Regards,
Abu Naser
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--
Jeff Hammond
jeff.science(a)gmail.com<mailto: jeff.science(a)gmail.com>
http: //jeffhammond.github.io/
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--_000_BLUPR0501MB2003EB7C1C3701600F398E4997770BLUPR0501MB2003_
Content-Type: text/html; charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable
<html><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Diso-8859-=
1">
<style type=3D"text/css" style=3D"display: none;"><!-- P {margin-top:0;margi=
n-bottom: 0;} --></style>
</head>
<body dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt;color:#000000;font=
-family: Calibri,Helvetica,sans-serif;" dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt; color: rgb(0, 0,=
0); font-family: Calibri, Helvetica, sans-serif, "EmojiFont", &q=
uot;Apple Color Emoji", "Segoe UI Emoji", NotoColorEmoji, &q=
uot;Segoe UI Symbol", "Android Emoji", EmojiSymbols;" dir=3D=
"ltr">
<p style=3D"margin-top: 0;margin-bottom:0">Hello Min,</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">Thanks for the clarification.&nbs=
p; I will do the experiment.<br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<div id=3D"Signature">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p>Thanks.</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr style=3D"display: inline-block;width:98%" tabindex=3D"-1">
<div id=3D"divRplyFwdMsg" dir=3D"ltr"><font style=3D"font-size: 11pt" face=
=3D"Calibri, sans-serif" color=3D"#000000"><b>From: </b> Min Si <msi(a)anl.=
gov><br>
<b>Sent: </b> Wednesday, June 20, 2018 1:39:30 PM<br>
<b>To: </b> discuss(a)mpich.org<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3DWindows-1252">
<div style=3D"background-color: #FFFFFF">Hi Abu,<br>
<br>
I think Jeff means that you should run your experiment with more iterations=
in order to get a stable results.<br>
- Increase the iteration of for loop in each execution (I think osu benchma=
rk allows you to set it)<br>
- Run the experiments 10 or 100 times, and take the average and standard de=
viation.<br>
<br>
If you see a very small standard deviation (e.g., <=3D5%), then the tren=
d is stable and you might not see such gaps.<br>
<br>
Best regards,<br>
Min<br>
<div class=3D"x_moz-cite-prefix">On 2018/06/20 12: 14, Abu Naser wrote:<br>
</div>
<blockquote type=3D"cite">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; co=
lor: rgb(0, 0, 0); font-family: Calibri, Helvetica, sans-serif, "Emoji=
Font", "Apple Color Emoji", "Segoe UI Emoji", Noto=
ColorEmoji, "Segoe UI Symbol", "Android Emoji", EmojiSy=
mbols;">
<p style=3D"margin-top: 0; margin-bottom:0">Hello Jeff,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Yes, I am using a switch and oth=
er machines are also connected with that switch.
<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">If I remove other machines and j=
ust use my two node with the switch, then will it improve the performance b=
y 200 ~ 400 iterations?</p>
<p style=3D"margin-top: 0; margin-bottom:0">Meanwhile I will give a try with=
a single dedicated cable.
<span></span><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Thank you.<br>
</p>
<div id=3D"x_Signature">
<div id=3D"x_divtagdefaultwrapper" dir=3D"ltr" style=3D"">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr tabindex=3D"-1" style=3D"display: inline-block; width:98%">
<div id=3D"x_divRplyFwdMsg" dir=3D"ltr"><font style=3D"font-size: 11pt" face=
=3D"Calibri, sans-serif" color=3D"#000000"><b>From: </b> Jeff Hammond
<a class=3D"x_moz-txt-link-rfc2396E OWAAutoLink" href=3D"mailto: jeff.scienc=
e(a)gmail.com" id=3D"LPlnk983157" previewremoved=3D"true">
<jeff.science(a)gmail.com></a><br>
<b>Sent: </b> Wednesday, June 20, 2018 12:52:06 PM<br>
<b>To: </b> MPICH<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3Dutf-8">
<div>
<div dir=3D"ltr">Is the ethernet connection a single dedicated cable betwee=
n the two machines or are you running through a switch that handles other t=
raffic?
<div><br>
</div>
<div>My best guess is that this is noise and that you may be able to avoid =
it by running a very long time, e.g. 10000 iterations.</div>
<div><br>
</div>
<div>Jeff</div>
</div>
<div class=3D"x_x_gmail_extra"><br>
<div class=3D"x_x_gmail_quote">On Wed, Jun 20, 2018 at 6: 53 AM, Abu Naser <=
span dir=3D"ltr">
<<a href=3D"mailto: an16e(a)my.fsu.edu" target=3D"_blank" id=3D"LPlnk305789=
" class=3D"OWAAutoLink" previewremoved=3D"true">an16e(a)my.fsu.edu</a>></s=
pan> wrote: <br>
<blockquote class=3D"x_x_gmail_quote" style=3D"margin: 0 0 0 .8ex; border-le=
ft: 1px #ccc solid; padding-left:1ex">
<div dir=3D"ltr">
<div id=3D"x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr" style=
=3D"">
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Good day to all,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I had run point to point osu_lat=
ency test in two nodes for 200 times. Followings are the average time=
in microsecond for various size of the messages -</p>
<div>1KB 84.8514 us<br>
<span style=3D"color: rgb(255,0,0)">2KB 73.52535</span> us=
<br>
4KB 272.55275 us<br>
<span style=3D"color: rgb(255,0,0)">8KB 234.86385</span> u=
s<br>
16KB 288.88 us<br>
32KB 523.3725 us<br>
64KB 910.4025 us</div>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">From the above looks like, 2KB m=
essage has less latency than 1 KB and 8KB has less latency than 4KB.
<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I was looking for explanation of=
this behavior but did not get any.</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<ol style=3D"margin-bottom: 0px; margin-top:0px">
<li><span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span><span> is set to 128K=
B. So none of the above message size is using Rendezvous protocol. Is there=
any partition inside eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB -=
64KB)? If yes then what are the
boundaries for them? Can I log them with debug-event-logging? </span><br>
</li></ol>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Setup I am using:</p>
<p style=3D"margin-top: 0; margin-bottom:0">- two nodes has intel core i7, o=
ne with 16gb memory another one 8gb</p>
<p style=3D"margin-top: 0; margin-bottom:0">- mpich 3.2.1, configured and bu=
ild to use nemesis tcp</p>
<p style=3D"margin-top: 0; margin-bottom:0">- 1gb Ethernet connection</p>
<p style=3D"margin-top: 0; margin-bottom:0">- NFS is using for sharing<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">- osu_latency : uses MPI_Send an=
d MPI_Recv</p>
<p style=3D"margin-top: 0; margin-bottom:0">- <span>MPIR_CVAR_CH3_EAGER_MAX_=
MSG_<wbr>SIZE</span>=3D
<span>131072</span> (128KB)<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Can anyone help me on that? Than=
ks in advance.<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<div id=3D"x_x_m_6077755676379859201Signature">
<div id=3D"x_x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr" style=
=3D"">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
</div>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" id=3D"LPlnk816471" class=3D"OWAAutoLink" previewremoved=3D"true">discuss@=
mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank" id=3D"LPlnk624595" class=3D"OWAAutoLink" previewremo=
ved=3D"true">https: //lists.mpich.org/<wbr>mailman/listinfo/discuss</a><br>
<br>
</blockquote>
</div>
<br>
<br clear=3D"all">
<div><br>
</div>
-- <br>
<div class=3D"x_x_gmail_signature">Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank" id=3D"LPlnk3149=
93" class=3D"OWAAutoLink" previewremoved=3D"true">jeff.science(a)gmail.com</a=
><br>
<a href=3D"http: //jeffhammond.github.io/" target=3D"_blank" id=3D"LPlnk8614=
34" class=3D"OWAAutoLink" previewremoved=3D"true">http: //jeffhammond.github=
.io/</a></div>
</div>
</div>
<br>
<fieldset class=3D"x_mimeAttachmentHeader"></fieldset> <br>
<pre>_______________________________________________=0A=
discuss mailing list <a class=3D"x_moz-txt-link-abbreviated OWAAutoLink=
" href=3D"mailto: discuss(a)mpich.org" id=3D"LPlnk657371" previewremoved=3D"tr=
ue">discuss(a)mpich.org</a>=0A=
To manage subscription options or unsubscribe: =0A=
<a class=3D"x_moz-txt-link-freetext OWAAutoLink" href=3D"https: //lists.mpic=
h.org/mailman/listinfo/discuss" id=3D"LPlnk669988" previewremoved=3D"true">=
https: //lists.mpich.org/mailman/listinfo/discuss</a>=0A=
</pre>
</blockquote>
<br>
</div>
</div>
</body>
</html>
--_000_BLUPR0501MB2003EB7C1C3701600F398E4997770BLUPR0501MB2003_--
--===============6609414382460160217==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============6609414382460160217==--
Message-ID: <sanitized-2243(a)migration.local>
1
0
as less latency than 4KB.
I was looking for explanation of this behavior but did not get any.
1. MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE is set to 128KB. So none of the abov=
e message size is using Rendezvous protocol. Is there any partition inside =
eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB)? If yes then wh=
at are the boundaries for them? Can I log them with debug-event-logging?
Setup I am using:
- two nodes has intel core i7, one with 16gb memory another one 8gb
- mpich 3.2.1, configured and build to use nemesis tcp
- 1gb Ethernet connection
- NFS is using for sharing
- osu_latency: uses MPI_Send and MPI_Recv
- MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE=3D 131072 (128KB)
Can anyone help me on that? Thanks in advance.
Best Regards,
Abu Naser
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--
Jeff Hammond
jeff.science(a)gmail.com<mailto: jeff.science(a)gmail.com>
http: //jeffhammond.github.io/
--_000_BLUPR0501MB200383DFBCA0721E71DDAC0E97770BLUPR0501MB2003_
Content-Type: text/html; charset="us-ascii"
Content-Transfer-Encoding: quoted-printable
<html><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dus-ascii"=
>
<style type=3D"text/css" style=3D"display: none;"><!-- P {margin-top:0;margi=
n-bottom: 0;} --></style>
</head>
<body dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" style=3D"font-size: 12pt;color:#000000;font=
-family: Calibri,Helvetica,sans-serif;" dir=3D"ltr">
<p style=3D"margin-top: 0;margin-bottom:0">Hello Jeff,</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">Yes, I am using a switch and othe=
r machines are also connected with that switch.
<br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">If I remove other machines and ju=
st use my two node with the switch, then will it improve the performance by=
200 ~ 400 iterations?</p>
<p style=3D"margin-top: 0;margin-bottom:0">Meanwhile I will give a try with =
a single dedicated cable.
<span></span><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0;margin-bottom:0">Thank you.<br>
</p>
<div id=3D"Signature">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
<hr style=3D"display: inline-block;width:98%" tabindex=3D"-1">
<div id=3D"divRplyFwdMsg" dir=3D"ltr"><font face=3D"Calibri, sans-serif" st=
yle=3D"font-size: 11pt" color=3D"#000000"><b>From:</b> Jeff Hammond <jeff=
.science(a)gmail.com><br>
<b>Sent: </b> Wednesday, June 20, 2018 12:52:06 PM<br>
<b>To: </b> MPICH<br>
<b>Subject: </b> Re: [mpich-discuss] osu_latency test: why 8KB takes less ti=
me than 4KB and 2KB takes less time than 1KB?</font>
<div> </div>
</div>
<meta content=3D"text/html; charset=3Dutf-8">
<div>
<div dir=3D"ltr">Is the ethernet connection a single dedicated cable betwee=
n the two machines or are you running through a switch that handles other t=
raffic?
<div><br>
</div>
<div>My best guess is that this is noise and that you may be able to avoid =
it by running a very long time, e.g. 10000 iterations.</div>
<div><br>
</div>
<div>Jeff</div>
</div>
<div class=3D"x_gmail_extra"><br>
<div class=3D"x_gmail_quote">On Wed, Jun 20, 2018 at 6: 53 AM, Abu Naser <sp=
an dir=3D"ltr">
<<a href=3D"mailto: an16e(a)my.fsu.edu" target=3D"_blank">an16e(a)my.fsu.edu<=
/a>></span> wrote: <br>
<blockquote class=3D"x_gmail_quote" style=3D"margin: 0 0 0 .8ex; border-left=
: 1px #ccc solid; padding-left:1ex">
<div dir=3D"ltr">
<div id=3D"x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr" style=
=3D"font-size: 12pt; color:rgb(0,0,0); font-family:Calibri,Helvetica,sans-se=
rif,"EmojiFont","Apple Color Emoji","Segoe UI Emoj=
i",NotoColorEmoji,"Segoe UI Symbol","Android Emoji"=
;,EmojiSymbols">
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Good day to all,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I had run point to point osu_lat=
ency test in two nodes for 200 times. Followings are the average time=
in microsecond for various size of the messages -</p>
<p style=3D"margin-top: 0; margin-bottom:0"></p>
<div>1KB 84.8514 us<br>
<span style=3D"color: rgb(255,0,0)">2KB 73.52535</span> us=
<br>
4KB 272.55275 us<br>
<span style=3D"color: rgb(255,0,0)">8KB 234.86385</span> u=
s<br>
16KB 288.88 us<br>
32KB 523.3725 us<br>
64KB 910.4025 us</div>
<p></p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">From the above looks like, 2KB m=
essage has less latency than 1 KB and 8KB has less latency than 4KB.
<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I was looking for explanation of=
this behavior but did not get any.</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<ol style=3D"margin-bottom: 0px; margin-top:0px">
<li><span>MPIR_CVAR_CH3_EAGER_MAX_MSG_<wbr>SIZE</span><span> is set to 128K=
B. So none of the above message size is using Rendezvous protocol. Is there=
any partition inside eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB -=
64KB)? If yes then what are the
boundaries for them? Can I log them with debug-event-logging? </span><br>
</li></ol>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Setup I am using:</p>
<p style=3D"margin-top: 0; margin-bottom:0">- two nodes has intel core i7, o=
ne with 16gb memory another one 8gb</p>
<p style=3D"margin-top: 0; margin-bottom:0">- mpich 3.2.1, configured and bu=
ild to use nemesis tcp</p>
<p style=3D"margin-top: 0; margin-bottom:0">- 1gb Ethernet connection</p>
<p style=3D"margin-top: 0; margin-bottom:0">- NFS is using for sharing<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">- osu_latency : uses MPI_Send an=
d MPI_Recv</p>
<p style=3D"margin-top: 0; margin-bottom:0">- <span>MPIR_CVAR_CH3_EAGER_MAX_=
MSG_<wbr>SIZE</span>=3D
<span>131072</span> (128KB)<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Can anyone help me on that? Than=
ks in advance.<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<div id=3D"x_m_6077755676379859201Signature">
<div id=3D"x_m_6077755676379859201divtagdefaultwrapper" dir=3D"ltr" style=
=3D"font-size: 12pt; color:rgb(0,0,0); font-family:Calibri,Helvetica,sans-se=
rif,"EmojiFont","Apple Color Emoji","Segoe UI Emoj=
i",NotoColorEmoji,"Segoe UI Symbol","Android Emoji"=
;,EmojiSymbols">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
</div>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/<wbr>mailman/listinfo/discus=
s</a><br>
<br>
</blockquote>
</div>
<br>
<br clear=3D"all">
<div><br>
</div>
-- <br>
<div class=3D"x_gmail_signature">Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank">jeff.science@gm=
ail.com</a><br>
<a href=3D"http: //jeffhammond.github.io/" target=3D"_blank">http://jeffhamm=
ond.github.io/</a></div>
</div>
</div>
</body>
</html>
--_000_BLUPR0501MB200383DFBCA0721E71DDAC0E97770BLUPR0501MB2003_--
--===============3147691773054401677==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============3147691773054401677==--
Message-ID: <sanitized-2213(a)migration.local>
1
0
as less latency than 4KB.
I was looking for explanation of this behavior but did not get any.
1. MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE is set to 128KB. So none of the abov=
e message size is using Rendezvous protocol. Is there any partition inside =
eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB)? If yes then wh=
at are the boundaries for them? Can I log them with debug-event-logging?
Setup I am using:
- two nodes has intel core i7, one with 16gb memory another one 8gb
- mpich 3.2.1, configured and build to use nemesis tcp
- 1gb Ethernet connection
- NFS is using for sharing
- osu_latency: uses MPI_Send and MPI_Recv
- MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE=3D 131072 (128KB)
Can anyone help me on that? Thanks in advance.
Best Regards,
Abu Naser
--_000_BLUPR0501MB2003829D702157A57187438D97710BLUPR0501MB2003_
Content-Type: text/html; charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable
<html><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Diso-8859-=
1">
<style type=3D"text/css" style=3D"display: none;"><!-- P {margin-top:0;margi=
n-bottom: 0;} --></style>
</head>
<body dir=3D"ltr">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Good day to all,</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I had run point to point osu_lat=
ency test in two nodes for 200 times. Followings are the average time=
in microsecond for various size of the messages -</p>
<p style=3D"margin-top: 0; margin-bottom:0"></p>
<div>1KB 84.8514 us<br>
<span style=3D"color: rgb(255,0,0)">2KB 73.52535</span> us=
<br>
4KB 272.55275 us<br>
<span style=3D"color: rgb(255,0,0)">8KB 234.86385</span> u=
s<br>
16KB 288.88 us<br>
32KB 523.3725 us<br>
64KB 910.4025 us</div>
<p></p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">From the above looks like, 2KB m=
essage has less latency than 1 KB and 8KB has less latency than 4KB.
<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">I was looking for explanation of=
this behavior but did not get any.</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<ol style=3D"margin-bottom: 0px; margin-top:0px">
<li><span>MPIR_CVAR_CH3_EAGER_MAX_MSG_SIZE</span><span> is set to 128KB. So=
none of the above message size is using Rendezvous protocol. Is there any =
partition inside eager protocol (e.g. 0 - 512 bytes, 1KB - 8KB, 16KB - 64KB=
)? If yes then what are the boundaries
for them? Can I log them with debug-event-logging? </span><br>
</li></ol>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Setup I am using:</p>
<p style=3D"margin-top: 0; margin-bottom:0">- two nodes has intel core i7, o=
ne with 16gb memory another one 8gb</p>
<p style=3D"margin-top: 0; margin-bottom:0">- mpich 3.2.1, configured and bu=
ild to use nemesis tcp</p>
<p style=3D"margin-top: 0; margin-bottom:0">- 1gb Ethernet connection</p>
<p style=3D"margin-top: 0; margin-bottom:0">- NFS is using for sharing<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">- osu_latency : uses MPI_Send an=
d MPI_Recv</p>
<p style=3D"margin-top: 0; margin-bottom:0">- <span>MPIR_CVAR_CH3_EAGER_MAX_=
MSG_SIZE</span>=3D
<span>131072</span> (128KB)<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0">Can anyone help me on that? Than=
ks in advance.<br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<p style=3D"margin-top: 0; margin-bottom:0"><br>
</p>
<div id=3D"Signature">
<div id=3D"divtagdefaultwrapper" dir=3D"ltr" style=3D"font-size: 12pt; color=
: rgb(0,0,0); font-family:Calibri,Helvetica,sans-serif,"EmojiFont"=
,"Apple Color Emoji","Segoe UI Emoji",NotoColorEmoji,&q=
uot;Segoe UI Symbol","Android Emoji",EmojiSymbols">
<p><br>
</p>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Best Regards,</span></p>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<div align=3D"left"><span style=3D"font-size: 11pt; font-family:Calibri,Helv=
etica,sans-serif"></span></div>
<span style=3D"font-family: Calibri,Helvetica,sans-serif; font-size:10pt"></=
span>
<p align=3D"left"><span style=3D"font-size: 10pt; font-family:Calibri,Helvet=
ica,sans-serif">Abu Naser</span><br>
</p>
</div>
</div>
</div>
</body>
</html>
--_000_BLUPR0501MB2003829D702157A57187438D97710BLUPR0501MB2003_--
--===============5933607345979184983==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============5933607345979184983==--
Message-ID: <sanitized-2109(a)migration.local>
1
0
loads and stores in the MPI_Win_lock_all epochs using MPI_Fetch_and_op (see=
attached files).<br>
<br>
This version behaves very similar to the original code and also fails from =
time to time. Putting a sleep into the acquire busy loop (usleep(100)) will=
make the code "much more robust" (I hack, I know, but indicating=
some underlying race condition?!). Let me know if you see any problems in =
the way I am using MPI_Fetch_and_op in a busy loop. Flushing or syncing is =
not necessary in this case, right?<br>
<br>
All work is done with export MPIR_CVAR_ASYNC_PROGRESS=3D1 on mpich-3.2 and =
mpich-3.3a2<br>
<br>
On Wed, Mar 8, 2017 at 4: 21 PM, Halim Amer <<a href=3D"mailto:aamer@anl.=
gov" target=3D"_blank">aamer(a)anl.gov</a>> wrote: <br>
I cannot claim that I thoroughly verified the correctness of that code, so =
take it with a grain of salt. Please keep in mind that it is a test code fr=
om a tutorial book; those codes are meant for learning purposes not for dep=
loyment.<br>
<br>
If your goal is to have a high performance RMA lock, I suggest you to look =
into the recent HPDC'16 paper: "High-Performance Distributed RMA Locks=
".<br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a><br>
<br>
On 3/8/17 3: 06 AM, Ask Jakobsen wrote:<br>
You are absolutely correct, Halim. Removing the test lmem[nextRank] =3D=3D =
-1<br>
in release fixes the problem. Great work. Now I will try to understand why<=
br>
you are right. I hope the authors of the book will credit you for<br>
discovering the bug.<br>
<br>
So in conclusion you need to remove the above mentioned test AND enable<br>
asynchronous progression using the environment variable<br>
MPIR_CVAR_ASYNC_PROGRESS=3D1 in MPICH (BTW I still can't get the code to wo=
rk<br>
in openmpi).<br>
<br>
On Tue, Mar 7, 2017 at 5: 37 PM, Halim Amer <<a href=3D"mailto:aamer@anl.=
gov" target=3D"_blank">aamer(a)anl.gov</a>> wrote: <br>
<br>
detect that another process is being or already enqueued in the MCS<br>
queue.<br>
<br>
Actually the problem occurs only when the waiting process already enqueued<=
br>
itself, i.e., the accumulate operation on the nextRank field succeeded.<br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a> <<a href=3D"http: //www.mcs.anl.gov/%7Eaam=
er" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/%7Eaam<wbr>=
er</a>><br>
<br>
<br>
On 3/7/17 10: 29 AM, Halim Amer wrote:<br>
<br>
In the Release protocol, try removing this test: <br>
<br>
if (lmem[nextRank] =3D=3D -1) {<br>
If-Block;<br>
}<br>
<br>
but keep the If-Block.<br>
<br>
The hang occurs because the process releasing the MCS lock fails to<br>
detect that another process is being or already enqueued in the MCS queue.<=
br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a> <<a href=3D"http: //www.mcs.anl.gov/%7Eaam=
er" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/%7Eaam<wbr>=
er</a>><br>
<br>
<br>
On 3/7/17 6: 43 AM, Ask Jakobsen wrote:<br>
<br>
Thanks, Halim. I have now enabled asynchronous progress in MPICH (can't<br>
find something similar in openmpi) and now all ranks acquire the lock and<b=
r>
the program finish as expected. However if I put a while(1) loop<br>
around the<br>
acquire-release code in main.c it will fail again at random and go<br>
into an<br>
infinite loop. The simple unfair lock does not have this problem.<br>
<br>
On Tue, Mar 7, 2017 at 12: 44 AM, Halim Amer <<a href=3D"mailto:aamer@anl=
.gov" target=3D"_blank">aamer(a)anl.gov</a>> wrote: <br>
<br>
My understanding is that this code assumes asynchronous progress.<br>
An example of when the processes hang is as follows: <br>
<br>
1) P0 Finishes MCSLockAcquire()<br>
2) P1 is busy waiting in MCSLockAcquire() at<br>
do {<br>
MPI_Win_sync(win);<br>
} while (lmem[blocked] =3D=3D 1);<br>
3) P0 executes MCSLockRelease()<br>
4) P0 waits on MPI_Win_lock_all() inside MCSLockRlease()<br>
<br>
Hang!<br>
<br>
For P1 to get out of the loop, P0 has to get out of<br>
MPI_Win_lock_all() and<br>
executes its Compare_and_swap().<br>
<br>
For P0 to get out MPI_Win_lock_all(), it needs an ACK from P1 that it<br>
got<br>
the lock.<br>
<br>
P1 does not make communication progress because MPI_Win_sync is not<br>
required to do so. It only synchronizes private and public copies.<br>
<br>
For this hang to disappear, one can either trigger progress manually by<br>
using heavy-duty synchronization calls instead of Win_sync (e.g.,<br>
Win_unlock_all + Win_lock_all), or enable asynchronous progress.<br>
<br>
To enable asynchronous progress in MPICH, set the<br>
MPIR_CVAR_ASYNC_PROGRESS<br>
env var to 1.<br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a> <<a href=3D"http: //www.mcs.anl.gov/%7Eaam=
er" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/%7Eaam<wbr>=
er</a>> <<br>
<a href=3D"http: //www.mcs.anl.gov/%7Eaamer" rel=3D"noreferrer" target=3D"_b=
lank">http: //www.mcs.anl.gov/%7Eaame<wbr>r</a>><br>
<br>
<br>
On 3/6/17 1: 11 PM, Ask Jakobsen wrote:<br>
<br>
I am testing on x86_64 platform.<br>
<br>
I have tried to built both the mpich and the mcs lock code with -O0 to<br>
avoid agressive optimization. After your suggestion I have also<br>
tried to<br>
make volatile int *pblocked pointing to lmem[blocked] in the<br>
MCSLockAcquire<br>
function and volatile int *pnextrank pointing to lmem[nextRank] in<br>
MCSLockRelease, but it does not appear to make a difference.<br>
<br>
On suggestion from Richard Warren I have also tried building the code<br>
using<br>
openmpi-2.0.2 without any luck (however it appears to acquire the<br>
lock a<br>
couple of extra times before failing) which I find troubling.<br>
<br>
I think I will give up using local load/stores and will see if I can<br>
figure<br>
out if rewrite using MPI calls like MPI_Fetch_and_op as you suggest.<=
br>
Thanks for your help.<br>
<br>
On Mon, Mar 6, 2017 at 7: 20 PM, Jeff Hammond <<a href=3D"mailto:jeff.sci=
ence(a)gmail.com" target=3D"_blank">jeff.science(a)gmail.com</a>><br>
wrote: <br>
<br>
What processor architecture are you testing?<br>
<br>
<br>
Maybe set lmem to volatile or read it with MPI_Fetch_and_op rather<br>
than a<br>
load. MPI_Win_sync cannot prevent the compiler from caching *lmem<br>
in a<br>
register.<br>
<br>
Jeff<br>
<br>
On Sat, Mar 4, 2017 at 12: 30 AM, Ask Jakobsen <<a href=3D"mailto:afj@qey=
e-labs.com" target=3D"_blank">afj(a)qeye-labs.com</a>><br>
wrote: <br>
<br>
Hi,<br>
<br>
<br>
I have downloaded the source code for the MCS lock from the excellent<br>
book "Using Advanced MPI" from <a href=3D"http: //www.mcs.anl.gov/=
researc" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/resear=
c</a><br>
h/projects/mpi/usingmpi/exampl<wbr>es-advmpi/rma2/mcs-lock.c<br>
<br>
I have made a very simple piece of test code for testing the MCS lock<br>
but<br>
it works at random and often never escapes the busy loops in the<br>
acquire<br>
and release functions (see attached source code). The code appears<br>
semantically correct to my eyes.<br>
<br>
#include <stdio.h><br>
#include <mpi.h><br>
#include "mcs-lock.h"<br>
<br>
int main(int argc, char *argv[])<br>
{<br>
MPI_Win win;<br>
MPI_Init( &argc, &argv );<br>
<br>
MCSLockInit(MPI_COMM_WORLD, &win);<br>
<br>
int rank, size;<br>
MPI_Comm_rank(MPI_COMM_WORLD, &rank);<br>
MPI_Comm_size(MPI_COMM_WORLD, &size);<br>
<br>
printf("rank: %d, size: %d\n", rank, size);<br>
<br>
<br>
MCSLockAcquire(win);<br>
printf("rank %d aquired lock\n", rank); fflush(=
stdout);<br>
MCSLockRelease(win);<br>
<br>
<br>
MPI_Win_free(&win);<br>
MPI_Finalize();<br>
return 0;<br>
}<br>
<br>
<br>
I have tested on several hardware platforms and mpich-3.2 and<br>
mpich-3.3a2<br>
but with no luck.<br>
<br>
It appears that the MPI_Win_Sync are not "refreshing" the local<b=
r>
data or<br>
I<br>
have a bug I can't spot.<br>
<br>
A simple unfair lock like <a href=3D"http: //www.mcs.anl.gov/researc" rel=3D=
"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/researc</a><br>
h/projects/mpi/usingmpi/exampl<wbr>es-advmpi/rma2/ga_mutex1.c works<br>
perfectly.<br>
<br>
Best regards, Ask Jakobsen<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
--<br>
Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank">jeff.science@gm=
ail.com</a><br>
<a href=3D"http: //jeffhammond.github.io/" rel=3D"noreferrer" target=3D"_bla=
nk">http: //jeffhammond.github.io/</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<main.c><mcs-lock-fop.c><mcs-l<wbr>ock.h>________________=
________<wbr>_______________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
</blockquote>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
--<br>
Ask Jakobsen<br>
R&D<br>
<br>
Qeye Labs<br>
Lers=C3=B8 Parkall=C3=A9 107<br>
2100 Copenhagen =C3=98<br>
Denmark<br>
<br>
mobile: <a href=3D"tel:%2B45%202834%206936" value=3D"+4528346936" targe=
t=3D"_blank">+45 2834 6936</a><br>
email: afj(a)Qeye-Labs.com<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
</blockquote>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
</blockquote>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
</blockquote>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a></div></div></blockquote></div><br><br clear=3D"all"><div><br></div>--=
<br><div class=3D"m_-9117170903988053540gmail_signature" data-smartmail=3D=
"gmail_signature"><div dir=3D"ltr"><div><div dir=3D"ltr"><font size=3D"1"><=
b>Ask Jakobsen</b><br>R&D<br><br><span style=3D"color: rgb(255,153,102)"=
>Q</span>eye Labs<br>Lers=C3=B8 Parkall=C3=A9 107<br>2100 Copenhagen =C3=98=
<br>Denmark<br><br>mobile: <a href=3D"tel:+45%2028%2034%2069%2036" val=
ue=3D"+4528346936" target=3D"_blank">+45 2834 6936</a><br>email: af=
j(a)Qeye-Labs.com<br></font></div></div></div></div>
</div>
</div></div></blockquote></div><br><br clear=3D"all"><div><br></div>-- <br>=
<div class=3D"gmail_signature" data-smartmail=3D"gmail_signature"><div dir=
=3D"ltr"><div><div dir=3D"ltr"><font size=3D"1"><b>Ask Jakobsen</b><br>R&am=
p;D<br><br><span style=3D"color: rgb(255,153,102)">Q</span>eye Labs<br>Lers=
=C3=B8 Parkall=C3=A9 107<br>2100 Copenhagen =C3=98 <br>Denmark<br><br>mobil=
e: +45 2834 6936<br>email: afj(a)Qeye-Labs.com<br></font></div></div></di=
v></div>
</div>
--001a113d3a84bc7a97054aa1fd98--
--===============6869846251208948639==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============6869846251208948639==--
Message-ID: <sanitized-2104(a)migration.local>
1
0
e algorithm. i.e. the problem may not lie with the inefficiency of your com=
munication between threads but just that your algorithm is not keeping the =
processor busy enough for a large number of threads. The only way that you =
will know for sure whether this is a comms issue or an algorithmic one is t=
o use a profiling tool, such as Vampir or Paraver.
With the profiling result you will be able to determine whether you need to=
make algorithmic changes in your bulk processing and enhance comms as per =
Huiwei's notes.
I would be interested to know what your profiling shows.
Regards, bob
On Thu, Oct 23, 2014 at 12: 02 AM, Qiguo Jing <qjing(a)trinityconsultants.com<=
mailto: qjing(a)trinityconsultants.com>> wrote:
Hi All,
We have a parallel program running on a cluster. We recently found a case,=
which decreases the CPU usage and increase the run-time when increases Nod=
es. Below is the results table.
The particular run requires a lot of data communication between nodes.
Any thoughts about this phenomena? Or is there any way we can improve the =
CPU usage when using higher number of nodes?
Average CPU Usage (%)
Number of Nodes
Number of Threads/Node
100
1
8
92
2
8
50
3
8
40
4
8
35
5
8
30
6
8
25
7
8
20
8
8
20
8
4
Thanks!
_________________________________________________________________________
The information transmitted is intended only for the person or entity to
which it is addressed and may contain confidential and/or privileged
material. Any review, retransmission, dissemination or other use of, or
taking of any action in reliance upon, this information by persons or
entities other than the intended recipient is prohibited. If you received
this in error, please contact the sender and delete the material from any
computer.
_________________________________________________________________________
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
_________________________________________________________________________
The information transmitted is intended only for the person or entity to
which it is addressed and may contain confidential and/or privileged
material. Any review, retransmission, dissemination or other use of, or
taking of any action in reliance upon, this information by persons or
entities other than the intended recipient is prohibited. If you received
this in error, please contact the sender and delete the material from any
computer.
_________________________________________________________________________
_________________________________________________________________________
The information transmitted is intended only for the person or entity to
which it is addressed and may contain confidential and/or privileged
material. Any review, retransmission, dissemination or other use of, or
taking of any action in reliance upon, this information by persons or
entities other than the intended recipient is prohibited. If you received
this in error, please contact the sender and delete the material from any
computer.
_________________________________________________________________________
_______________________________________________
discuss mailing list discuss(a)mpich.org<mailto: discuss(a)mpich.org>
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--
/ Ruben FAELENS
+32 494 06 72 59
--=20
_________________________________________________________________________
The information transmitted is intended only for the person or entity to
which it is addressed and may contain confidential and/or privileged
material. Any review, retransmission, dissemination or other use of, or
taking of any action in reliance upon, this information by persons or
entities other than the intended recipient is prohibited. If you received
this in error, please contact the sender and delete the material from any
computer.
_________________________________________________________________________
--_000_9D079B25EA4E3149B260846F07AE80754C5944A3TCIEXCH03ustrin_
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
<html xmlns: v=3D"urn:schemas-microsoft-com:vml" xmlns:o=3D"urn:schemas-micr=
osoft-com: office:office" xmlns:w=3D"urn:schemas-microsoft-com:office:word" =
xmlns: m=3D"http://schemas.microsoft.com/office/2004/12/omml" xmlns=3D"http:=
//www.w3.org/TR/REC-html40"><head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dutf-8">
<meta name=3D"Generator" content=3D"Microsoft Word 15 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
{font-family: =E5=AE=8B=E4=BD=93;
panose-1:2 1 6 0 3 1 1 1 1 1;}
@font-face
{font-family: "Cambria Math";
panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
{font-family: Calibri;
panose-1:2 15 5 2 2 2 4 3 2 4;}
@font-face
{font-family: "\@=E5=AE=8B=E4=BD=93";
panose-1:2 1 6 0 3 1 1 1 1 1;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
{margin: 0in;
margin-bottom:.0001pt;
font-size:12.0pt;
font-family:"Times New Roman",serif;}
a: link, span.MsoHyperlink
{mso-style-priority:99;
color:blue;
text-decoration:underline;}
a: visited, span.MsoHyperlinkFollowed
{mso-style-priority:99;
color:purple;
text-decoration:underline;}
span.EmailStyle17
{mso-style-type: personal-reply;
font-family:"Calibri",sans-serif;
color:#1F497D;}
.MsoChpDefault
{mso-style-type: export-only;
font-family:"Calibri",sans-serif;}
@page WordSection1
{size: 8.5in 11.0in;
margin:1.0in 1.25in 1.0in 1.25in;}
div.WordSection1
{page: WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o: shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o: shapelayout v:ext=3D"edit">
<o: idmap v:ext=3D"edit" data=3D"1" />
</o: shapelayout></xml><![endif]-->
</head>
<body lang=3D"EN-US" link=3D"blue" vlink=3D"purple">
<div class=3D"WordSection1">
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D">Hi Ruben,<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D">You are right! My algorithm does what=
you described.
<o: p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D">I will record the timestamp for each =
thread and every event. Thanks for your suggestions!
<o: p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D">Qiguo<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"font-size: 11.0pt;font-family:"Ca=
libri",sans-serif;color: #1F497D"><o:p> </o:p></span></p>
<p class=3D"MsoNormal"><b><span style=3D"font-size: 11.0pt;font-family:"=
;Calibri",sans-serif">From: </span></b><span style=3D"font-size:11.0pt;=
font-family: "Calibri",sans-serif"> parasietje(a)gmail.com [mailto:p=
arasietje(a)gmail.com]
<b>On Behalf Of </b>Ruben Faelens<br>
<b>Sent: </b> Thursday, October 23, 2014 10:58 AM<br>
<b>To: </b> discuss(a)mpich.org<br>
<b>Subject: </b> Re: [mpich-discuss] CPU usage versus Nodes, Threads<o:p></o=
: p></span></p>
<p class=3D"MsoNormal"><o: p> </o:p></p>
<div>
<p class=3D"MsoNormal">Hi Qiguo,<o: p></o:p></p>
<div>
<p class=3D"MsoNormal"><o: p> </o:p></p>
</div>
<div>
<p class=3D"MsoNormal">You should try to collect performance statistics, es=
pecially regarding your specific nodes and what they are doing at every mom=
ent in time.<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">If I understand correctly, your algorithm does the f=
ollowing: <o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- Master thread: read in data, split it up into piec=
es, transfer pieces to slaves<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- Slave thread: do calculation, transfer data back t=
o master<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- Master: recombine data, do calculation, split data=
back up, transfer pieces to slaves<o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- etc...<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal"><o: p> </o:p></p>
</div>
<div>
<p class=3D"MsoNormal">The reason you do not see linear performance scaling=
could be due to the following: <o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- The master thread recombining and splitting the da=
ta set may be responsible for a large part of the work (and therefore is th=
e bottleneck)<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- Work is not divided equally. A significant part of=
the time is spent waiting on one slave node who has a more difficult probl=
em (takes longer) than the rest.<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- There is a common dataset. I/O takes a larger part=
of the time when more slave nodes are used.<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal"><o: p> </o:p></p>
</div>
<div>
<p class=3D"MsoNormal">The only way to know for sure, is to simply generate=
a log file that shows the time every process starts and ends a specific pr=
ocess step. Output the time when <o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- the slave starts receiving data<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- starts calculation<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- starts sending results back<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal">- starts waiting for his next piece of data<o: p></o:=
p></p>
</div>
<div>
<p class=3D"MsoNormal">This will clearly show you what each node is doing a=
t each moment in time, and should identify the bottleneck.<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal"><o: p> </o:p></p>
</div>
<div>
<p class=3D"MsoNormal">/ Ruben<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal"><o: p> </o:p></p>
</div>
</div>
<div>
<p class=3D"MsoNormal"><o: p> </o:p></p>
<div>
<p class=3D"MsoNormal">On Thu, Oct 23, 2014 at 5: 27 PM, Qiguo Jing <<a h=
ref=3D"mailto: qjing(a)trinityconsultants.com" target=3D"_blank">qjing@trinity=
consultants.com</a>> wrote: <o:p></o:p></p>
<blockquote style=3D"border: none;border-left:solid #CCCCCC 1.0pt;padding:0i=
n 0in 0in 6.0pt;margin-left: 4.8pt;margin-right:0in">
<div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Hi Bob,</span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Thanks for your suggestions. Here are m=
ore tests. We actually have three clusters.</span><o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Cluster 1 and 2: 8 nodes, (2 Processors, 4 co=
res/processor, no HT =E2=80=93 Total 8 Threads)/node</span><o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Cluster 3:  =
; 8 nodes, (1 Processors, 4 cores /proc=
essor, HT =E2=80=93 Total 8 Threads )/node</span><o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">We also have a standalone machine: 2 processo=
rs, 6 cores/processor, HT =E2=80=93 total 24 threads.</span><o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">For one particular case:
</span><o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Cluster 1 and 2 take 48 min to finish with 8 nodes,=
8 threads/node, 60% CPU usage; 53 min to finish
with 3 nodes, 8 threads/node, 90% CPU usage;</span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Cluster 3 takes 227 min to finish with 8 nodes, 8 t=
hreads/node, 20% CPU usage; 207 min to finish with
3 nodes, 8 threads/node, 50% CPU usage;</span><o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Standalone machine takes 82 min to finish with 24 t=
hreads, 100% CPU usage.</span><o: p></o:p></p>
<div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">It looks like with 24 threads, they should be prett=
y busy? Could the above phenomena be a hardware
issue more than software?</span><o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D">Qiguo</span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><span style=3D"font-size:11.0pt;font-family:"Calibri",sa=
ns-serif;color: #1F497D"> </span><o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><b><span style=3D"font-size:11.0pt;font-family:"Calibri"=
,sans-serif">From: </span></b><span style=3D"font-size:11.0pt;font-family:&q=
uot;Calibri",sans-serif"> Bob Ilgner [<a href=3D"mailto: bobilgner@gmai=
l.com" target=3D"_blank">mailto: bobilgner(a)gmail.com</a>]
<br>
<b>Sent: </b> Thursday, October 23, 2014 1:11 AM<br>
<b>To: </b> <a href=3D"mailto:[email protected]" target=3D"_blank">discuss@m=
pich.org</a><br>
<b>Subject: </b> Re: [mpich-discuss] CPU usage versus Nodes, Threads</span><=
o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">Hi Qiguo,<o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">From the results table it looks as if you are using a computa=
tionally sparse algorithm. i.e. the problem may not lie with the inefficien=
cy of your communication between threads
but just that your algorithm is not keeping the processor busy enough for =
a large number of threads. The only way that you will know for sure whether=
this is a comms issue or an algorithmic one is to use a profiling tool, su=
ch as Vampir or Paraver.<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">With the profiling result you will be able to determine whether yo=
u need to make algorithmic changes in your bulk processing and enhance comm=
s as per Huiwei's notes.<o: p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">I would be interested to know what your profiling shows.<o:p></o:p=
></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">Regards, bob<o:p></o:p></p>
</div>
</div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">On Thu, Oct 23, 2014 at 12:02 AM, Qiguo Jing <<a href=3D"mailto=
: qjing(a)trinityconsultants.com" target=3D"_blank">qjing(a)trinityconsultants.c=
om</a>> wrote: <o:p></o:p></p>
<blockquote style=3D"border: none;border-left:solid #CCCCCC 1.0pt;padding:0i=
n 0in 0in 6.0pt;margin-left: 4.8pt;margin-top:5.0pt;margin-right:0in;margin-=
bottom: 5.0pt">
<div>
<div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">Hi All,<o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">We have a parallel program running on a cluster. We recently=
found a case, which decreases the CPU usage and increase the run-time when=
increases Nodes. Below is the results
table.<o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">The particular run requires a lot of data communication between no=
des.
<o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">Any thoughts about this phenomena? Or is there any way we ca=
n improve the CPU usage when using higher number of nodes?
<o: p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<table class=3D"MsoNormalTable" border=3D"0" cellspacing=3D"0" cellpadding=
=3D"0" width=3D"441" style=3D"width: 331.0pt;border-collapse:collapse">
<tbody>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">Average CPU Usage (%)</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r: solid windowtext 1.0pt;border-left:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">Number of Nodes</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er: solid windowtext 1.0pt;border-left:none;padding:0in 5.4pt 0in 5.4pt;heig=
ht: 15.0pt;border-color:currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">Number of Threads/Node</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">100</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">1</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">92</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">2</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">50</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">3</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">40</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">4</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">35</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">5</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">30</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">6</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">25</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">7</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">20</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
</tr>
<tr style=3D"height: 15.0pt">
<td width=3D"153" nowrap=3D"" valign=3D"bottom" style=3D"width: 115.0pt;bord=
er: solid windowtext 1.0pt;border-top:none;padding:0in 5.4pt 0in 5.4pt;heigh=
t: 15.0pt;border-color:currentColor windowtext windowtext">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">20</span><o:p></o:p></p>
</td>
<td width=3D"119" nowrap=3D"" valign=3D"bottom" style=3D"width: 89.0pt;borde=
r-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-rig=
ht: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border-=
color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">8</span><o:p></o:p></p>
</td>
<td width=3D"169" nowrap=3D"" valign=3D"bottom" style=3D"width: 127.0pt;bord=
er-top: none;border-left:none;border-bottom:solid windowtext 1.0pt;border-ri=
ght: solid windowtext 1.0pt;padding:0in 5.4pt 0in 5.4pt;height:15.0pt;border=
-color: currentColor windowtext windowtext currentColor">
<p class=3D"MsoNormal" align=3D"center" style=3D"mso-margin-top-alt: auto;ms=
o-margin-bottom-alt: auto;text-align:center">
<span style=3D"color: black">4</span><o:p></o:p></p>
</td>
</tr>
</tbody>
</table>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto">Thanks!<o:p></o:p></p>
</div>
</div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><br>
_________________________________________________________________________<b=
r>
<br>
The information transmitted is intended only for the person or entity to<br=
>
which it is addressed and may contain confidential and/or privileged<br>
material. Any review, retransmission, dissemination or other use of, or<br>
taking of any action in reliance upon, this information by persons or<br>
entities other than the intended recipient is prohibited. If you received<b=
r>
this in error, please contact the sender and delete the material from any<b=
r>
computer.<br>
_________________________________________________________________________<b=
r>
<br>
_______________________________________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" target=3D"_bla=
nk">https: //lists.mpich.org/mailman/listinfo/discuss</a><o:p></o:p></p>
</blockquote>
</div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"> <o:p></o:p></p>
</div>
<p class=3D"MsoNormal" style=3D"mso-margin-top-alt: auto;mso-margin-bottom-a=
lt: auto"><br>
_________________________________________________________________________<b=
r>
<br>
The information transmitted is intended only for the person or entity to<br=
>
which it is addressed and may contain confidential and/or privileged<br>
material. Any review, retransmission, dissemination or other use of, or<br>
taking of any action in reliance upon, this information by persons or<br>
entities other than the intended recipient is prohibited. If you received<b=
r>
this in error, please contact the sender and delete the material from any<b=
r>
computer.<br>
_________________________________________________________________________<o=
: p></o:p></p>
</div>
</div>
</div>
</div>
<div>
<div>
<p class=3D"MsoNormal"><br>
_________________________________________________________________________<b=
r>
<br>
The information transmitted is intended only for the person or entity to<br=
>
which it is addressed and may contain confidential and/or privileged<br>
material. Any review, retransmission, dissemination or other use of, or<br>
taking of any action in reliance upon, this information by persons or<br>
entities other than the intended recipient is prohibited. If you received<b=
r>
this in error, please contact the sender and delete the material from any<b=
r>
computer.<br>
_________________________________________________________________________<o=
: p></o:p></p>
</div>
</div>
<p class=3D"MsoNormal"><br>
_______________________________________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" target=3D"_bla=
nk">https: //lists.mpich.org/mailman/listinfo/discuss</a><o:p></o:p></p>
</blockquote>
</div>
<p class=3D"MsoNormal"><br>
<br clear=3D"all">
<o: p></o:p></p>
<div>
<p class=3D"MsoNormal"><o: p> </o:p></p>
</div>
<p class=3D"MsoNormal">-- <br>
/ Ruben FAELENS<o: p></o:p></p>
<div>
<p class=3D"MsoNormal">+32 494 06 72 59<o: p></o:p></p>
</div>
</div>
</div>
</body>
</html>
<br>
______________________________<wbr>______________________________<wbr>_____=
________<br><br>The information transmitted is intended only for the person=
or entity to<br>which it is addressed and may contain confidential and/or =
privileged<br>material. Any review, retransmission, dissemination or other=
use of, or<br>taking of any action in reliance upon, this information by p=
ersons or<br>entities other than the intended recipient is prohibited. If=
you received<br>this in error, please contact the sender and delete the ma=
terial from any<br>computer.<br>______________________________<wbr>________=
______________________<wbr>_____________<br>=
--_000_9D079B25EA4E3149B260846F07AE80754C5944A3TCIEXCH03ustrin_--
--===============3243062315966390760==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============3243062315966390760==--
Message-ID: <sanitized-2083(a)migration.local>
1
0
loads and stores in the MPI_Win_lock_all epochs using MPI_Fetch_and_op (see=
attached files).<br>
<br>
This version behaves very similar to the original code and also fails from =
time to time. Putting a sleep into the acquire busy loop (usleep(100)) will=
make the code "much more robust" (I hack, I know, but indicating=
some underlying race condition?!). Let me know if you see any problems in =
the way I am using MPI_Fetch_and_op in a busy loop. Flushing or syncing is =
not necessary in this case, right?<br>
<br>
All work is done with export MPIR_CVAR_ASYNC_PROGRESS=3D1 on mpich-3.2 and =
mpich-3.3a2<br>
<br>
On Wed, Mar 8, 2017 at 4: 21 PM, Halim Amer <<a href=3D"mailto:aamer@anl.=
gov" target=3D"_blank">aamer(a)anl.gov</a>> wrote: <br>
I cannot claim that I thoroughly verified the correctness of that code, so =
take it with a grain of salt. Please keep in mind that it is a test code fr=
om a tutorial book; those codes are meant for learning purposes not for dep=
loyment.<br>
<br>
If your goal is to have a high performance RMA lock, I suggest you to look =
into the recent HPDC'16 paper: "High-Performance Distributed RMA Locks=
".<br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a><br>
<br>
On 3/8/17 3: 06 AM, Ask Jakobsen wrote:<br>
You are absolutely correct, Halim. Removing the test lmem[nextRank] =3D=3D =
-1<br>
in release fixes the problem. Great work. Now I will try to understand why<=
br>
you are right. I hope the authors of the book will credit you for<br>
discovering the bug.<br>
<br>
So in conclusion you need to remove the above mentioned test AND enable<br>
asynchronous progression using the environment variable<br>
MPIR_CVAR_ASYNC_PROGRESS=3D1 in MPICH (BTW I still can't get the code to wo=
rk<br>
in openmpi).<br>
<br>
On Tue, Mar 7, 2017 at 5: 37 PM, Halim Amer <<a href=3D"mailto:aamer@anl.=
gov" target=3D"_blank">aamer(a)anl.gov</a>> wrote: <br>
<br>
detect that another process is being or already enqueued in the MCS<br>
queue.<br>
<br>
Actually the problem occurs only when the waiting process already enqueued<=
br>
itself, i.e., the accumulate operation on the nextRank field succeeded.<br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a> <<a href=3D"http: //www.mcs.anl.gov/%7Eaam=
er" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/%7Eaam<wbr>=
er</a>><br>
<br>
<br>
On 3/7/17 10: 29 AM, Halim Amer wrote:<br>
<br>
In the Release protocol, try removing this test: <br>
<br>
if (lmem[nextRank] =3D=3D -1) {<br>
If-Block;<br>
}<br>
<br>
but keep the If-Block.<br>
<br>
The hang occurs because the process releasing the MCS lock fails to<br>
detect that another process is being or already enqueued in the MCS queue.<=
br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a> <<a href=3D"http: //www.mcs.anl.gov/%7Eaam=
er" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/%7Eaam<wbr>=
er</a>><br>
<br>
<br>
On 3/7/17 6: 43 AM, Ask Jakobsen wrote:<br>
<br>
Thanks, Halim. I have now enabled asynchronous progress in MPICH (can't<br>
find something similar in openmpi) and now all ranks acquire the lock and<b=
r>
the program finish as expected. However if I put a while(1) loop<br>
around the<br>
acquire-release code in main.c it will fail again at random and go<br>
into an<br>
infinite loop. The simple unfair lock does not have this problem.<br>
<br>
On Tue, Mar 7, 2017 at 12: 44 AM, Halim Amer <<a href=3D"mailto:aamer@anl=
.gov" target=3D"_blank">aamer(a)anl.gov</a>> wrote: <br>
<br>
My understanding is that this code assumes asynchronous progress.<br>
An example of when the processes hang is as follows: <br>
<br>
1) P0 Finishes MCSLockAcquire()<br>
2) P1 is busy waiting in MCSLockAcquire() at<br>
do {<br>
MPI_Win_sync(win);<br>
} while (lmem[blocked] =3D=3D 1);<br>
3) P0 executes MCSLockRelease()<br>
4) P0 waits on MPI_Win_lock_all() inside MCSLockRlease()<br>
<br>
Hang!<br>
<br>
For P1 to get out of the loop, P0 has to get out of<br>
MPI_Win_lock_all() and<br>
executes its Compare_and_swap().<br>
<br>
For P0 to get out MPI_Win_lock_all(), it needs an ACK from P1 that it<br>
got<br>
the lock.<br>
<br>
P1 does not make communication progress because MPI_Win_sync is not<br>
required to do so. It only synchronizes private and public copies.<br>
<br>
For this hang to disappear, one can either trigger progress manually by<br>
using heavy-duty synchronization calls instead of Win_sync (e.g.,<br>
Win_unlock_all + Win_lock_all), or enable asynchronous progress.<br>
<br>
To enable asynchronous progress in MPICH, set the<br>
MPIR_CVAR_ASYNC_PROGRESS<br>
env var to 1.<br>
<br>
Halim<br>
<a href=3D"http: //www.mcs.anl.gov/~aamer" rel=3D"noreferrer" target=3D"_bla=
nk">www.mcs.anl.gov/~aamer</a> <<a href=3D"http: //www.mcs.anl.gov/%7Eaam=
er" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/%7Eaam<wbr>=
er</a>> <<br>
<a href=3D"http: //www.mcs.anl.gov/%7Eaamer" rel=3D"noreferrer" target=3D"_b=
lank">http: //www.mcs.anl.gov/%7Eaame<wbr>r</a>><br>
<br>
<br>
On 3/6/17 1: 11 PM, Ask Jakobsen wrote:<br>
<br>
I am testing on x86_64 platform.<br>
<br>
I have tried to built both the mpich and the mcs lock code with -O0 to<br>
avoid agressive optimization. After your suggestion I have also<br>
tried to<br>
make volatile int *pblocked pointing to lmem[blocked] in the<br>
MCSLockAcquire<br>
function and volatile int *pnextrank pointing to lmem[nextRank] in<br>
MCSLockRelease, but it does not appear to make a difference.<br>
<br>
On suggestion from Richard Warren I have also tried building the code<br>
using<br>
openmpi-2.0.2 without any luck (however it appears to acquire the<br>
lock a<br>
couple of extra times before failing) which I find troubling.<br>
<br>
I think I will give up using local load/stores and will see if I can<br>
figure<br>
out if rewrite using MPI calls like MPI_Fetch_and_op as you suggest.<=
br>
Thanks for your help.<br>
<br>
On Mon, Mar 6, 2017 at 7: 20 PM, Jeff Hammond <<a href=3D"mailto:jeff.sci=
ence(a)gmail.com" target=3D"_blank">jeff.science(a)gmail.com</a>><br>
wrote: <br>
<br>
What processor architecture are you testing?<br>
<br>
<br>
Maybe set lmem to volatile or read it with MPI_Fetch_and_op rather<br>
than a<br>
load. MPI_Win_sync cannot prevent the compiler from caching *lmem<br>
in a<br>
register.<br>
<br>
Jeff<br>
<br>
On Sat, Mar 4, 2017 at 12: 30 AM, Ask Jakobsen <<a href=3D"mailto:afj@qey=
e-labs.com" target=3D"_blank">afj(a)qeye-labs.com</a>><br>
wrote: <br>
<br>
Hi,<br>
<br>
<br>
I have downloaded the source code for the MCS lock from the excellent<br>
book "Using Advanced MPI" from <a href=3D"http: //www.mcs.anl.gov/=
researc" rel=3D"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/resear=
c</a><br>
h/projects/mpi/usingmpi/exampl<wbr>es-advmpi/rma2/mcs-lock.c<br>
<br>
I have made a very simple piece of test code for testing the MCS lock<br>
but<br>
it works at random and often never escapes the busy loops in the<br>
acquire<br>
and release functions (see attached source code). The code appears<br>
semantically correct to my eyes.<br>
<br>
#include <stdio.h><br>
#include <mpi.h><br>
#include "mcs-lock.h"<br>
<br>
int main(int argc, char *argv[])<br>
{<br>
MPI_Win win;<br>
MPI_Init( &argc, &argv );<br>
<br>
MCSLockInit(MPI_COMM_WORLD, &win);<br>
<br>
int rank, size;<br>
MPI_Comm_rank(MPI_COMM_WORLD, &rank);<br>
MPI_Comm_size(MPI_COMM_WORLD, &size);<br>
<br>
printf("rank: %d, size: %d\n", rank, size);<br>
<br>
<br>
MCSLockAcquire(win);<br>
printf("rank %d aquired lock\n", rank); fflush(=
stdout);<br>
MCSLockRelease(win);<br>
<br>
<br>
MPI_Win_free(&win);<br>
MPI_Finalize();<br>
return 0;<br>
}<br>
<br>
<br>
I have tested on several hardware platforms and mpich-3.2 and<br>
mpich-3.3a2<br>
but with no luck.<br>
<br>
It appears that the MPI_Win_Sync are not "refreshing" the local<b=
r>
data or<br>
I<br>
have a bug I can't spot.<br>
<br>
A simple unfair lock like <a href=3D"http: //www.mcs.anl.gov/researc" rel=3D=
"noreferrer" target=3D"_blank">http: //www.mcs.anl.gov/researc</a><br>
h/projects/mpi/usingmpi/exampl<wbr>es-advmpi/rma2/ga_mutex1.c works<br>
perfectly.<br>
<br>
Best regards, Ask Jakobsen<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
--<br>
Jeff Hammond<br>
<a href=3D"mailto: jeff.science(a)gmail.com" target=3D"_blank">jeff.science@gm=
ail.com</a><br>
<a href=3D"http: //jeffhammond.github.io/" rel=3D"noreferrer" target=3D"_bla=
nk">http: //jeffhammond.github.io/</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<main.c><mcs-lock-fop.c><mcs-l<wbr>ock.h>________________=
________<wbr>_______________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
</blockquote>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
<br>
<br>
--<br>
Ask Jakobsen<br>
R&D<br>
<br>
Qeye Labs<br>
Lers=C3=B8 Parkall=C3=A9 107<br>
2100 Copenhagen =C3=98<br>
Denmark<br>
<br>
mobile: <a href=3D"tel:%2B45%202834%206936" value=3D"+4528346936" targe=
t=3D"_blank">+45 2834 6936</a><br>
email: afj(a)Qeye-Labs.com<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
</blockquote>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
</blockquote>
<br>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a><br>
<br>
</blockquote>
______________________________<wbr>_________________<br>
discuss mailing list <a href=3D"mailto: discuss(a)mpich.org=
" target=3D"_blank">discuss(a)mpich.org</a><br>
To manage subscription options or unsubscribe: <br>
<a href=3D"https: //lists.mpich.org/mailman/listinfo/discuss" rel=3D"norefer=
rer" target=3D"_blank">https: //lists.mpich.org/mailma<wbr>n/listinfo/discus=
s</a></div></div></blockquote></div><br><br clear=3D"all"><div><br></div>--=
<br><div class=3D"gmail_signature" data-smartmail=3D"gmail_signature"><div=
dir=3D"ltr"><div><div dir=3D"ltr"><font size=3D"1"><b>Ask Jakobsen</b><br>=
R&D<br><br><span style=3D"color: rgb(255,153,102)">Q</span>eye Labs<br>L=
ers=C3=B8 Parkall=C3=A9 107<br>2100 Copenhagen =C3=98 <br>Denmark<br><br>mo=
bile: +45 2834 6936<br>email: afj(a)Qeye-Labs.com<br></font></div></div><=
/div></div>
</div>
--94eb2c13e2e428ae25054aa11616--
--===============4174043369255602394==
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
discuss mailing list discuss(a)mpich.org
To manage subscription options or unsubscribe:
https: //lists.mpich.org/mailman/listinfo/discuss
--===============4174043369255602394==--
Message-ID: <sanitized-2078(a)migration.local>
1
0