Asg
Threads by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
February 2012
- 15 participants
- 91 discussions
aesop Repository branch, master, updated. a3ac9a69b1bb6965f2727af7deef3175877c7548
by noreply@mcs.anl.gov 28 Feb '12
by noreply@mcs.anl.gov 28 Feb '12
28 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via a3ac9a69b1bb6965f2727af7deef3175877c7548 (commit)
from d0bad468a57892628ba200f2d2c1ac3a464ad831 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit a3ac9a69b1bb6965f2727af7deef3175877c7548
Author: Kevin Harms <harms(a)alcf.anl.gov>
Date: Tue Feb 28 14:18:29 2012 -0600
Minor edits to aug_introduction.txt
-----------------------------------------------------------------------
Summary of changes:
doc/aug_introduction.txt | 9 ++++++---
1 files changed, 6 insertions(+), 3 deletions(-)
Diff of changes:
diff --git a/doc/aug_introduction.txt b/doc/aug_introduction.txt
index 6fd3f36..3ea60d9 100644
--- a/doc/aug_introduction.txt
+++ b/doc/aug_introduction.txt
@@ -123,9 +123,12 @@ If mpich is installed, there is no need to install OpenPA.
The Open Portable Atomic Library can be downloaded from the following
location: http://trac.mcs.anl.gov/projects/openpa.
Configure using the included `configure` script, optionally specifying where
-the library needs to be installed (using the `--prefix` option). If installing
-in a non-standard location, make sure to use the `--with-openpa`
+the library needs to be installed (using the `--prefix` option).
+[NOTE]
+=====
+If installing OpenPA in a non-standard location, make sure to use the `--with-openpa` argument when configuring Aesop.
+=====
Obtaining the Aesop distribution
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
@@ -148,7 +151,7 @@ repository might contain bugs or fail to build. If that is the case, please
let us know by opening a ticket at http://trac.mcs.anl.gov/projects/aesop.
=========
-In order to install Aesop is is necessary to first clone the repository.
+In order to install Aesop it is necessary to first clone the repository.
The aesop source code repository is public and can be cloned by
anonymous users:
hooks/post-receive
--
aesop Repository
1
0
aesop Repository branch, master, updated. d0bad468a57892628ba200f2d2c1ac3a464ad831
by noreply@mcs.anl.gov 28 Feb '12
by noreply@mcs.anl.gov 28 Feb '12
28 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via d0bad468a57892628ba200f2d2c1ac3a464ad831 (commit)
from 158db93595548dbf3fb2245df5fa947b17aebba3 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit d0bad468a57892628ba200f2d2c1ac3a464ad831
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Tue Feb 28 12:17:08 2012 -0500
make graphviz figure smaller
-----------------------------------------------------------------------
Summary of changes:
doc/aug_introduction.txt | 2 +-
1 files changed, 1 insertions(+), 1 deletions(-)
Diff of changes:
diff --git a/doc/aug_introduction.txt b/doc/aug_introduction.txt
index c6eb37a..6fd3f36 100644
--- a/doc/aug_introduction.txt
+++ b/doc/aug_introduction.txt
@@ -289,7 +289,7 @@ Using the Aesop Source-To-Source Translator
File Dependencies
~~~~~~~~~~~~~~~~~~
-[graphviz]
+["graphviz", width=300]
.Translation Dependencies
---------------
digraph G
hooks/post-receive
--
aesop Repository
1
0
aesop Repository branch, master, updated. 158db93595548dbf3fb2245df5fa947b17aebba3
by noreply@mcs.anl.gov 27 Feb '12
by noreply@mcs.anl.gov 27 Feb '12
27 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via 158db93595548dbf3fb2245df5fa947b17aebba3 (commit)
from 0555e006fec3cf568a2b686eea559b7a68a62e07 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 158db93595548dbf3fb2245df5fa947b17aebba3
Author: Kevin Harms <harms(a)alcf.anl.gov>
Date: Mon Feb 27 23:56:23 2012 -0600
Minor edits to performance paper, fix typos and grammar.
-----------------------------------------------------------------------
Summary of changes:
doc/aesop-performance.txt | 89 +++++++++++++++++++++------------------------
1 files changed, 41 insertions(+), 48 deletions(-)
Diff of changes:
diff --git a/doc/aesop-performance.txt b/doc/aesop-performance.txt
index 942fd2e..a25e782 100644
--- a/doc/aesop-performance.txt
+++ b/doc/aesop-performance.txt
@@ -17,7 +17,7 @@ The remainder of this document is organized as follows.
case study which is then used to evaluate the performance of Aesop
relative to more traditional server architectures in terms of performance
(<<runtime-perf>>), memory usage (<<sec-memory>>), and productivity
-(<<sec-productivity>>). <<sec-overhead-analyis>> provides a more detailed
+(<<sec-productivity>>). <<sec-overhead-analysis>> provides a more detailed
breakdown of specific sources of Aesop overhead, while
<<sec-compile>> concludes by discussing Aesop compile-time code translation
performance.
@@ -37,13 +37,13 @@ of the test server and client are given in the following subsections.
Each of the example servers used for comparison in this document implement
an identical request protocol and are therefore evaluated using the same
client test harness. The client is a basic C program that uses TCP sockets to send messages to
-the server. It use MPI to coordinate processes and generate a highly
+the server. It uses MPI to coordinate processes and generate a highly
concurrent workload.
The client will execute in a loop generating a specified number of
operations to the server. The general flow is that each client process sends a request
to the server that contains an optional payload. The client then waits for
-the server to send an acknowledgment that also contain an optional payload.
+the server to send an acknowledgment that also contains an optional payload.
==== Request Types
@@ -250,18 +250,11 @@ the number of clients from
server nodes were used used when making comparisons across implementations
for a given workload. Fusion is a shared resource and may experience
increased network contention at times. We therefore executed each
-test case five times and cycled between server implementations in a
-round-robin fashion to minimize the possibility of any given server
+test case five times to minimize the possibility of any given server
execution being unfairly penalized by external contention. We show the
-median result in all graphs unless
-otherwise noted. The server daemon was restarted on each
-iteration for two reasons. First, this approach allowed us to capture memory statistics independently for each
-run and insure that results were not affected by resources left over from
-previous runs. Secondly, it allowed us to cycle between servers as
-described above without any risk of one daemon interfering with the
-performance of another daemon.
-
-On fusion we determined that we could use 16 client processes per physical node. Using
+median result of the five iterations in all graphs unless otherwise noted.
+
+On Fusion we determined that we could use 16 client processes per physical node. Using
more clients per node caused the bottleneck of the test to become the client
nodes instead of the server. Therefore, for the largest scale tests shown
in this study we utilized 65 total nodes. One node acted as the server,
@@ -278,7 +271,7 @@ performed with TCP/IP sockets.
===== Read
The read test had clients each issue 16 requests asking for 4 KiB from
-disk. Each client specifies a unique file to be read on on each request. All
+disk. Each client specifies a unique file to be read on each request. All
clients specify unique files. The files are first generated by a script
that runs before the read test starts. The script generates files for every
client in the local storage of the server.
@@ -424,7 +417,7 @@ image::fig/write-time.png[]
[[fig-readtime]]
image::fig/read-time.png[]
-In <<fig-readtime>> and <<fig-writetime>> we see that the event server is
+In <<fig-writetime>> and <<fig-readtime>> we see that the event server is
the only fair server. However, it achieves this fairness by sacrificing
overall throughput. During our initial experimentation, aesop was configured
to be a fair system as it favored executing all disk and network operations
@@ -438,15 +431,15 @@ for it to favor "hot" requests as opposed to FIFO ordering.
Interestingly, we see that the thread-per-client, thread-per-client-nb, and
thread-per-op servers are more fair on the write test than on the read test.
-.Fastest and Slowest Total Client Runtime for Read-Null Test
-[[fig-readnulltime]]
-image::fig/read-null-time.png[]
-
.Fastest and Slowest Total Client Runtime for Write-Null Test
[[fig-writenulltime]]
image::fig/write-null-time.png[]
-The <<fig-readnulltime>> and <<fig-writenulltime>> show similar fairness between
+.Fastest and Slowest Total Client Runtime for Read-Null Test
+[[fig-readnulltime]]
+image::fig/read-null-time.png[]
+
+The <<fig-writenulltime>> and <<fig-readnulltime>> show similar fairness between
all server types except thread-pool in the case of the write-null test case.
The thread-pool server is the only one that shows a notable difference in
fairness due to its scheduling approach.
@@ -462,13 +455,13 @@ image::fig/write-lat.png[]
image::fig/read-lat.png[]
The last client metric is an examination of per operation latency. The
-client measures the latency of each individual request and the computes the minimum and
+client measures the latency of each individual request and then computes the minimum and
maximum latency, the first quartile latency and third quartile latency. These
metrics are graphed using a box and whiskers plot. The box represents the
first and third quartiles and the whiskers are the minimum and maximum
values. The following results are for the 1024 client size selected from
-the same iteration as the maximum runtime graphs. <<fig-readlat>> and
-<<fig-writelat>> show the aesop offers vary comparative latency performance
+the same iteration as the maximum runtime graphs. <<fig-writelat>> and
+<<fig-readlat>> show aesop offers vary comparative latency performance
as the other configurations and only notably thread-per-op and event are
significantly worse. An interesting observation is that the overall fairness
across clients shown in the previous section does not appear to correlate in
@@ -482,7 +475,7 @@ image::fig/write-null-lat.png[]
[[fig-readnulllat]]
image::fig/read-null-lat.png[]
-In <<fig-readnulllat>> and <<fig-writenulllat>> we see again that aesop has
+In <<fig-writenulllat>> and <<fig-readnulllat>> we see again that aesop has
very good latency metrics compared to the other servers. In this case the
thread-per-op and event are noticeably worse than the other server
implementations.
@@ -491,7 +484,7 @@ implementations.
The runtime analysis shows that aesop is competitive with the most
optimal alternate server implementations at the 1024 client size.
-In the _read_ test, aesop is 4.5% slower that the fastest implementation,
+In the _read_ test, aesop is 4.5% slower than the fastest implementation,
thread-pool. Aesop was the fastest implementation in the _write_ test.
The thread-per-client implementation was fastest in the _read null_ and
_write null_ tests. Here aesop is significantly slower at around 25-30%,
@@ -505,28 +498,28 @@ or SSM, we believe that it would perform as well as a thread-per-client implemen
Another interesting factor was that there was no one fastest server
implementation. Each workload presents a different challenge and building
a single tuned server is difficult. A key advantage of aesop is the ability
-to change the underlying implementation/turning for resources or the aesop
-runtime without changing the aesop server source. During our investigation
+to change the underlying implementation or turning for the aesop
+runtime without changing the server code. During our investigation
we experimented with different underlying thread models for the network and
file resources of aesop. The aesop resources can be configured at runtime
to use the different thread models or tuning options. With the knowledge
-that building a specific server implementation that is optimal for a generic
-workload, aesop becomes powerful because these types of changes can be made
-without ever changing the server source.
+that building a specific server implementation, which is optimal for a generic
+workload, is extremely difficult, aesop becomes powerful because these types of changes can be made
+without changing the server source.
[[sec-memory]]
== Runtime Memory Efficiency
Another aspect of the overall performance is the memory efficiency of each
server implementation. In this section we compare aesop to the other server implementations
-as we did for the runtime performance.
+using the same runtime performance experitment.
=== Experiment
The memory utilization of each server implementation was captured during
the performance experiments detailed in <<runtime-perf>>. We recorded
-the VmHWM stat from the server when the client test was completed. The
-VmHWM stat is a Linux-specific metric that represents the peak resident
+the *VmHWM* stat from the server when the client test was completed. The
+*VmHWM* stat is a Linux-specific metric that represents the peak resident
set size (RSS) of an executable, where RSS corresponds to the amount
of paged-in memory used by the executable.
@@ -542,7 +535,7 @@ image::fig/write-mem.png[]
[[fig-readmem]]
image::fig/read-mem.png[]
-In <<fig-readmem>> and <<fig-writemem>> we see that thread-pool limits the
+In <<fig-writemem>> and <<fig-readmem>> we see that thread-pool limits the
memory usage as the client work load increases because the thread-pool
by design limits the number of requests that can be in progress at once. The
other server implementations scale as the number of clients increase.
@@ -555,7 +548,7 @@ image::fig/write-null-mem.png[]
[[fig-readnullmem]]
image::fig/read-null-mem.png[]
-<<fig-readnullmem>> and <<fig-writenullmem>> show a similar result as the
+<<fig-writenullmem>> and <<fig-readnullmem>> show a similar result as the
disk I/O tests. In this case the event server also produces favorable
results (in addition to the thread pool server) because it is the only
implementation which does not spawn any additional threads to handle
@@ -577,7 +570,7 @@ to be a more relevant metric in practice.
The core design element of aesop is to make programming of a concurrent
server easier, so the trade off for memory and runtime performance should
-be worth it. To evaluate this we examine the code complexity of each of
+be worth it. To evaluate this we examine the code complexity of each of the
server implementations.
.Implementation complexity analysis.
@@ -639,7 +632,7 @@ Mod. CC, qualitatively it is significantly more challenging to develop.
== Sources of Aesop overhead
As we've seen aesop compares favorably to other concurrency models but loses
-some performance compared to best case hand tuned version. Here we examine
+some performance compared to the best case hand tuned version. Here we examine
the overheads associated with aesop. The first item to examine is the
cost associated with an aesop blocking call.
@@ -668,7 +661,7 @@ code. While aesop does not create any threads, many of the aesop resources
internally use threads. Therefore, the aesop compiler has to ensure that the
emitted code is multi-thread safe. For example, in the following code,
two pbranches might be executing concurrently using different threads.
-Since these threads could both by modifying the pwait state simultaneously,
+Since these threads could both be modifying the pwait state simultaneously,
aesop
uses locks to serialize their access. Locks, and other synchronization
primitives account for most of the remaining performance difference between
@@ -734,14 +727,14 @@ The results for these are shown in the +tcmalloc+ column.
The program used to obtain these results is in the repository:
+tests/blocking-overhead.ae+.
-From these results we see that an Aesop __blocking function is significantly
+From these results we see that an Aesop `__blocking` function is significantly
slower than a basic C function. Test case 4 illustrates that almost all of
this overhead is a result of the memory management and synchronization
performed by Aesop in order to enable efficient concurrency.
It is also important to note that Aesop is a superset of the C language, and
normal C functions (with the associated low overhead) can still be used in
-an Aesop program. The __blocking functions are most appropriate for
+an Aesop program. The `__blocking` functions are most appropriate for
functions that perform blocking device or resource operations, while
standard C functions are better suited for computationally intense inner-loop routines.
@@ -781,13 +774,13 @@ consume memory until the pbranch returns.
[[sec-compile]]
== Compile Time Performance
-Currently aesop imposes some overhead when compiling aesop source. This is
-demonstrated in a simple micro-benchmark. The test takes an existing aesop
-source fill with 4200 lines or source and is about 116 KB in size and compiles
-it to an object file. The same source file is then renamed to a .c file and
-four +#define+ are added which redefine the aesop keywords to nothing.
+Currently aesop imposes some overhead when compiling aesop source.
+In a simple micro-benchmark an existing aesop source file was compiled and
+then compiled again as straight C. The aesop source file contains 4043
+lines and is about 116 KB in size. The same source file is renamed to a .c
+file and four +#define+s are added which redefine the aesop keywords to nothing.
-The test is executed on a dual processor Intel Xeon E5620 running at 2.4 GHz
+The test was executed on a dual processor Intel Xeon E5620 running at 2.4 GHz
with 24 GB of RAM.
.Compile Time Comparison
@@ -799,7 +792,7 @@ with 24 GB of RAM.
| C | 1.22
|============================
-The performance penalty is significant but this only effects development. It
+The performance penalty is significant but only effects development. It
can be mitigated by constructing a Makefile that supports parallel make.
== Bibliography
hooks/post-receive
--
aesop Repository
1
0
aesop Repository branch, master, updated. 0555e006fec3cf568a2b686eea559b7a68a62e07
by noreply@mcs.anl.gov 27 Feb '12
by noreply@mcs.anl.gov 27 Feb '12
27 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via 0555e006fec3cf568a2b686eea559b7a68a62e07 (commit)
from 5a4048595bdf6d394e72b40bfa307555eefefd81 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 0555e006fec3cf568a2b686eea559b7a68a62e07
Author: Dries Kimpe <dkimpe(a)mcs.anl.gov>
Date: Mon Feb 27 23:24:42 2012 -0600
Improve document based on Rob's comments
-----------------------------------------------------------------------
Summary of changes:
doc/aug_external.txt | 17 +++++++++--------
doc/aug_introduction.txt | 4 ++--
doc/aug_language.txt | 14 +++++++++-----
doc/aug_standard.txt | 38 +++++++++++++++++++++++++++++---------
4 files changed, 49 insertions(+), 24 deletions(-)
Diff of changes:
diff --git a/doc/aug_external.txt b/doc/aug_external.txt
index 59ab2c2..b51fad5 100644
--- a/doc/aug_external.txt
+++ b/doc/aug_external.txt
@@ -14,7 +14,7 @@ function should be blocking or not.
Non-blocking functions
~~~~~~~~~~~~~~~~~~~~~~
-If the the external library function does not need to be blocking,
+If the external library function does not need to be blocking,
the function can be called directly. No special considerations need to be
made.
@@ -26,7 +26,7 @@ call them as in a regular C program.
Asynchronous Interfaces
~~~~~~~~~~~~~~~~~~~~~~~
-However, in some cases, the external library provides an explicit asynchronous
+In some cases, the external library provides an explicit asynchronous
interface. Consider the POSIX aio functions. The POSIX aio library supports
starting a non-blocking operation through `aio_read` or `aio_write`, with
notification either through a signal or by calling a user-specified callback
@@ -34,7 +34,7 @@ function.
While it is possible to directly use the asynchronous interface to the
library, doing so does not take advantage of the extra features aesop
-provides. For example, overlap (concurrency) has to be managed manually and
+provides. For example, overlap (concurrency) has to be managed manually, and
aesop will not be able to cancel ongoing operations.
In addition, in aesop, the functionality offered by the library would
@@ -190,17 +190,18 @@ directly integrate an external library into aesop.
Resources?
~~~~~~~~~~
-<<ref-whenblocking>> describes how to construct an aesop blocking function.
-However, it is also possible to create a blocking function _directly in C_.
-This is done by creating a new aesop _resource_.
+<<ref-whenblocking>> describes how to construct a blocking function in aesop.
+However, it is also possible for a pure C library to export an aesop blocking
+function, fully supporting features such as concurrent execution and
+cancellation. This is done by creating a new aesop _resource_.
In short, a resource integrates with aesop to export a number of regular and
_blocking_ functions, together with hooks aesop can call to cancel the
blocking calls issued by the resource, and to enable the resource to make
progress.
-Note that resources are implemented in C, and do not have access to any of the
-aesop extensions.
+NOTE: Resources are implemented entirely in C, and therefore cannot use any
+of the aesop extenions.
Defining Blocking Functions
~~~~~~~~~~~~~~~~~~~~~~~~~~~~
diff --git a/doc/aug_introduction.txt b/doc/aug_introduction.txt
index a80c83e..c6eb37a 100644
--- a/doc/aug_introduction.txt
+++ b/doc/aug_introduction.txt
@@ -18,7 +18,7 @@ Prerequisites
These are the requirements of the aesop source-to-source translator.
Executables compiled from aesop source code do not require any additional
packages, except for those required used by the program itself. For example,
-there is no need to have a haskell installed when running an executable
+there is no need to have haskell installed when running an executable
created using the aesop source-to-source translator.
====
@@ -67,7 +67,7 @@ Glasgow Haskell Compiler
[TIP]
Many linux distributions, such
as Ubuntu and Debian, contain this package. If your linux distribution has a
-pacakge for the Glasgow Haskell Compiler (typically named 'ghc'), we recommend
+package for the Glasgow Haskell Compiler (typically named 'ghc'), we recommend
installing ghc on your system using the distribution's package manager.
If your distribution does not include a package, or you do not want to install
the package system-wide, the compiler can be installed using the instructions
diff --git a/doc/aug_language.txt b/doc/aug_language.txt
index ad0af62..5c7a230 100644
--- a/doc/aug_language.txt
+++ b/doc/aug_language.txt
@@ -31,7 +31,7 @@ also valid aesop code.
.Implementation Detail
====
The current implementation of the aesop translator, while supporting many C99
-features, does not yet support all of the C99 functionality , in particular
+features, does not yet support all of the C99 functionality, in particular
those related to variable declarations in locations other than the beginning
of the function. These issues will be fixed in later releases.
====
@@ -45,7 +45,7 @@ it is legal to call that function (whether blocking or not) concurrently
from multiple threads.
Aesop programs exposing concurrency through aesop's language features (see
-<<ref-pbranch>> have to be thread-safe, as the implementation may choose to
+<<ref-pbranch>>) have to be thread-safe, as the implementation may choose to
use multiple threads to execute code whenever possible.
@@ -57,7 +57,7 @@ Blocking Functions
//===================================================================
//===================================================================
-Aesop extends the C language with a additional function type: blocking
+Aesop extends the C language with an additional function type: blocking
functions. Blocking functions support concurrent
execution, and are used by aesop to introduce concurrency to
the C language.
@@ -148,7 +148,7 @@ idle until the read completes.
Parallel Branches
------------------
-This section explores the concept of parallel branches (or /pbranch/).
+This section explores the concept of parallel branches (or _pbranch_).
Parallel Branches
@@ -189,6 +189,10 @@ In branch 2
After pwait
----
+NOTE: This is one possible output from the program. Since there is no
+synchronization between both branches, the output of line 7 and line 12 could
+appear in any order, including intertwined.
+
In this example, the second `printf` (line 12) executes without waiting for
the first timer call (line 8) to complete. The `aesop_timer` call is a
blocking call, so after starting the timer, execution continued in the next
@@ -396,7 +400,7 @@ Failing to do so will cause a compilation error.
The aesop main function
~~~~~~~~~~~~~~~~~~~~~~~
-The `main` function for an aesop program is identified using `aesop_set_main`.
+The `main` function for an aesop program is identified using `aesop_main_set`.
The named function must be a blocking function, and must take the same
arguments as the standard C `main` function.
diff --git a/doc/aug_standard.txt b/doc/aug_standard.txt
index 1a73dfd..3b0c841 100644
--- a/doc/aug_standard.txt
+++ b/doc/aug_standard.txt
@@ -49,14 +49,22 @@ information to `stderr`.
The following components can be traced:
-[horizontal]
-*blocking*:: Outputs when a blocking call is initiated and finished.
+[cols=">1,6", frame="none", grid="none"]
+|=====
-*cancel*:: Outputs information regarding cancellation of blocking calls.
+| *blocking* | Outputs when a blocking call is initiated and finished.
-*pbranch*:: Tracks the creation and completion of pbranches.
+| |
-The following call can be used to enable or disable tracing for a component.
+| *cancel* | Outputs information regarding cancellation of blocking calls.
+
+| |
+
+| *pbranch* | Tracks the creation and completion of pbranches.
+
+|=====
+
+The `aesop_set_debugging` call can be used to enable or disable tracing for a component.
----
int aesop_set_debugging (const char * what, int value);
@@ -119,6 +127,15 @@ Module initialization functions.
----
+int aesocket_prepare (int fd);
+----
+
+Before a socket can be used with the aesocket functions, it needs some tuning
+(for example switching it to non-blocking mode). The `aesocket_prepare`
+function takes care of this. This function might change the
+blocking/non-blocking nature of the given descriptor.
+
+----
__blocking int aesocket_accept(
int sockfd,
struct sockaddr *addr,
@@ -134,7 +151,8 @@ The function returns `AE_SUCCESS` when a new connection is accepted,
`AE_ERR_CANCELED` if the call was canceled and `AE_ERR_OTHER` if a system call
error occurred (returned in `*err`) occurred.
-
+NOTE: It is not necessary to call `aesocket_prepare` on the descriptors
+returned by this function.
-----
__blocking int aesocket_read(
@@ -209,7 +227,7 @@ These are aesop versions of the regular `pwrite`, `pread`, `fsync`,
====
These functions are currently implemented using a thread which calls the
regular POSIX I/O function. These functions cannot be cancelled once the
-operation started.
+operation has started.
====
@@ -320,8 +338,10 @@ thread to become available. The function returns `AE_ERR_CANCELLED` when the
search was cancelled, or `AE_SUCCESS` in when execution successfully switched
to one of the group threads.
-The thread module is mainly used to provide a blocking version of a regular C
-function. See <<ref-thread>> for more details and an example.
+The thread module is mainly used to create an aesop blocking function (which
+supports concurrent execution) from a long lived C function which would
+otherwise force the calling thread to become idle.
+See <<ref-thread>> for more details and an example.
hooks/post-receive
--
aesop Repository
1
0
aesop Repository branch, master, updated. 5a4048595bdf6d394e72b40bfa307555eefefd81
by noreply@mcs.anl.gov 27 Feb '12
by noreply@mcs.anl.gov 27 Feb '12
27 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via 5a4048595bdf6d394e72b40bfa307555eefefd81 (commit)
from 66ac30aaaf1f4b1174b29c5bb89a19cc3491784d (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 5a4048595bdf6d394e72b40bfa307555eefefd81
Author: Dries Kimpe <dkimpe(a)mcs.anl.gov>
Date: Mon Feb 27 18:16:49 2012 -0600
Finish last chapter
-----------------------------------------------------------------------
Summary of changes:
doc/aug_external.txt | 477 +++++++++++++++++++++++++++++++++++++++-----------
doc/aug_standard.txt | 52 +-----
2 files changed, 386 insertions(+), 143 deletions(-)
Diff of changes:
diff --git a/doc/aug_external.txt b/doc/aug_external.txt
index 3a04257..59ab2c2 100644
--- a/doc/aug_external.txt
+++ b/doc/aug_external.txt
@@ -1,192 +1,470 @@
Interfacing with external C libraries
=====================================
+
+[[ref-aesop-techniques]]
+Aesop Techniques
+----------------
+
+This chapter describes how to call existing C code from an aesop program.
+The way to do this depends on if the aesop protoype of the external
+function should be blocking or not.
+
[[ref-interface-non-blocking]]
Non-blocking functions
-----------------------
+~~~~~~~~~~~~~~~~~~~~~~
+
+If the the external library function does not need to be blocking,
+the function can be called directly. No special considerations need to be
+made.
+
+As an example, consider the `math.h` header. All functions provided by this
+header are cpu-bound, meaning there is generally no reason to make them
+blocking in aesop. To call these functions, simply include the header and
+call them as in a regular C program.
+
+Asynchronous Interfaces
+~~~~~~~~~~~~~~~~~~~~~~~
+
+However, in some cases, the external library provides an explicit asynchronous
+interface. Consider the POSIX aio functions. The POSIX aio library supports
+starting a non-blocking operation through `aio_read` or `aio_write`, with
+notification either through a signal or by calling a user-specified callback
+function.
+
+While it is possible to directly use the asynchronous interface to the
+library, doing so does not take advantage of the extra features aesop
+provides. For example, overlap (concurrency) has to be managed manually and
+aesop will not be able to cancel ongoing operations.
+
+In addition, in aesop, the functionality offered by the library would
+typically be exposed through blocking functions, as I/O depends on external
+events (i.e. disk, network) and is not cpu bound.
+
+Therefore, it is recommended to create a single blocking aesop function
+combining the asynchronous operation and its completion or cancellation.
+There are a number of ways to do this. This section describes using the
+ResourceBuilder module (<<ref-resourcebuilder>>). An alternative technique is
+discussed in <<ref-resource>>.
-Non-blocking functions can be called directly.
+[[ref-resourcebuilder]]
+Building Blocking Functions using ResourceBuilder
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+This section discusses how to make a blocking aesop write function on top of
+the POSIX AIO asynchronous I/O functions.
+
+
+----
+ static void aesop_write_complete (union sigval)
+ {
+ rb_slot_complete ((rb_slot_t *) sigval.sival_ptr);
+ }
-If the library provides an asynchronous interface, even though
-Currently, there are two ways to do this.
+ __blocking int aesop_write (int fd, ... )
+ {
+ struct aiocb aio;
+ rb_slot_t slot;
-The first one is to use the _ResourceBuilder_ module (described below).
-The second one is to write an aesop resource for the library.
-ResourceBuilder
-^^^^^^^^^^^^^^^
+ aio.aio_fildes = fd;
+ ...
+ aio.aio_sigevent.sigev_notify = SIGEV_THREAD;
+ aio.aio_sigevent.sigev_value.sival_ptr = &slot;
+ aio.aio_sigevent.sigev_notify_function = &aesop_write_complete;
+ ...
+ rb_slot_initialize (&slot);
+
+ // start asynchronous operation
+ aio_write (&aio);
+
+ // wait for operation to complete
+ ret = rb_slot_capture (&slot);
+ if (ret != AE_SUCCESS)
+ {
+ // cancelled: handle appropriately
+ ...
+ aio_cancel (fd, &aio);
+ ...
+
+ return -EINTR;
+ }
+
+ rb_slot_destroy (&slot);
+
+ return ...
+ }
+----
+
+In this example, the resourcebuilder slot is used as a condition variable.
+The function starts the asynchronous I/O operation (which is a non-blocking
+call), and then waits for the slot to be signalled.
+
+The signalling happens by the completion of the I/O operation, which will call
+the `aio_write_complete` callback and complete the slot.
+
+Since the `rb_slot_capture` function is blocking, aesop can switch execution
+elsewhere until the slot is either completed (by the completion of the write
+operation) or until the call is cancelled (by a call to `ae_cancel_branches`),
+in which case the capture function will return with an error.
[[ref-interface-blocking]]
Blocking functions
--------------------
+~~~~~~~~~~~~~~~~~~
+C library functions which should really be blocking in aesop
+(to improve concurrency and ease of use) generally need some wrapper code to
+make them into blocking aesop functions. This section describes how this can
+be done using the 'thread' module.
+[[ref-thread]]
+Building Blocking Functions using the Thread module
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+Consider the POSIX open function:
-[[ref-thread]]
-Using the `thread` module
-^^^^^^^^^^^^^^^^^^^^^^^^^
+----
+int open (const char * pathname, int flags);
+----
+
+Since there is no asynchronous version of the `open` call, we cannot apply the
+ResourceBuilder technique (<<ref-resourcebuilder>>). The best way to deal with
+these functions is to create a thread to call the function without blocking
+the calling thread. The newly created thread will be
+idle while waiting for the operation to complete. However, aesop will use the
+parent thread to continue executing other code where possible.
+
+The thread module can simplify implementing this technique:
+
+----
+___blocking int aesop_open (const char * pathname, int flags)
+{
+ int ret;
+ ret = aethread_hint (open_group);
+ if (ret == AE_ERR_CANCELLED)
+ {
+ // operation was cancelled
+ return -EINTR;
+ }
+ return open (pathname, flags);
+}
+----
+
+In this example, `open_group` is a thread group which was created by an
+earlier call to `aethread_create_group`. As described in <<ref-thread>>,
+provided the `aethread_hint` call is successful, _in the current aesop
+implementation_, the code up to the following blocking call will execute in a
+thread borrowed from the thread group. As `aethread_hint` is a blocking
+function, the aesop thread calling `aesop_open` will switch to other work upon
+entering the `aethread_hint` call, making sure execution in other branches
+continues while the newly created thread will execute the `open` call and
+sleep until the operation completes.
[NOTE]
+.Implementation Detail
=====
-The BranchThreader module relies on the internals of the current
-source-to-source translator.
+The thread module relies on the internals of the current
+source-to-source translator. Its use and API might change in future aesop
+releases.
=====
//========================================================================
//========================================================================
-Resources
----------
+[[ref-resource]]
+Aesop Resources
+---------------
//========================================================================
//========================================================================
-In some cases, the techniques outlined in <<ref-interface-non-blocking>> and
-<<ref-interface-blocking>> cannot be used or do not offer sufficient control.
+In some cases, the techniques outlined in <<ref-aesop-techniques>>
+cannot be used or do not offer sufficient control.
This section provides details on how to write the low level glue code to
directly integrate an external library into aesop.
Resources?
~~~~~~~~~~
-Defining a new resource
-~~~~~~~~~~~~~~~~~~~~~~~
+<<ref-whenblocking>> describes how to construct an aesop blocking function.
+However, it is also possible to create a blocking function _directly in C_.
+This is done by creating a new aesop _resource_.
+
+In short, a resource integrates with aesop to export a number of regular and
+_blocking_ functions, together with hooks aesop can call to cancel the
+blocking calls issued by the resource, and to enable the resource to make
+progress.
+
+Note that resources are implemented in C, and do not have access to any of the
+aesop extensions.
+
+Defining Blocking Functions
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+A resource can define a new blocking call by using the `ae_define_post` macro.
+(This macro is defined in the `resource.h` header)
+
+For example:
----
-struct ae_resource
+ae_define_post (int, rb_slot_capture, rb_slot_t * slot)
{
- const char *resource_name;
- int (*test)(ae_op_id_t id, int ms_timeout);
- int (*poll_context)(ae_context_t context);
- int (*cancel)(ae_context_t ctx, ae_op_id_t id);
- int (*register_context)(ae_context_t context);
- int (*unregister_context)(ae_context_t context);
- struct ae_resource_config* config_array; /* terminated by entry with NULL name */
-};
+ ...
+}
----
-A resource is responsible for indicating it has work to do, by calling the
-`ae_resource_request_poll` function. After calling this function, aesop will
-schedule a call to the resource's `poll_context` function.
-The `ae_resource_request_poll` function is thread-safe and can safely be
-called from within a signal handler.
+The code above generates the following aesop prototype:
-[NOTE]
-====
-The _context_ concept is deprecated and will be removed in a future version.
-The `register_context` and `unregister_context` functions should be set to
-`NULL`. The `poll_context` function can safely ignore the `context` parameter.
-====
+----
+__blocking int rb_slot_capture (rb_slot_t * slot);
+----
-The `resource_name` member is the only mandatory field in this structure.
+The `ae_define_post` macro supports a variable number of arguments, enabling
+the creation of blocking functions with more than one parameter.
+The resource has to provide a matching aesop header file (`.hae') listing the
+public (aesop) prototypes of the functions it exports.
+The function needs to return a single integer argument. This argument can be
+one of the following:
+* `AE_SUCCESS`: the blocking function was successfully initiated.
+* `AE_IMMEDIATE_COMPLETION`: The blocking function already completed. This is
+ an optimization indicating aesop does not need to switch to another function
+ to continue execution. In effect, it makes the function behave as a
+ non-blocking function.
+* An aesop error code: this indicates that there was a problem starting the
+ blocking call.
-The new resource should be registered with aesop using the `ae_resource_register`
-function.
+The blocking function, assuming no immediate completion, needs to do the
+following:
+
+. Generate an _operation id_ (`ae_op_id_t` type), which can be used by aesop
+to test for completion and to identify the operation in case of cancellation.
+
+. Capture the context of the caller, so that execution can continue at this
+point once the blocking function completes.
+
+. Return from the function so that the thread can switch to other work (for
+example a concurrent branch) while waiting for the completion of the blocking
+call.
+
+Once the resource determines that the blocking function can complete, it will
+use the captured context to continue execution with the statement following
+the blocking call.
+
+Generating the Operation ID
+^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+The following function creates a new operation id:
----
-int ae_resource_register (struct ae_resource *resource, int *newid);
-void ae_resource_unregister (int resource_id);
+ae_op_id_t ae_id_gen (int resource_id, uintptr_t data);
----
+The operation id includes the identity of the resource that generated it
+(`resource_id`), and allows for private data (up to the size of a pointer) to
+be stored within the resource id. This is for internal use by the resource.
+The `resource_id` value is provided when registering the resource (see
+<<ref-register-resource>>).
-/////
+The private data can be retrieved using the `ae_id_lookup` function.
-What is a resource?
-A resource is a very thin shim layer that converts aesop blocking calls into async calls to some system resource or system library (like mpi, file access, ssm, etc.)
+----
+uintptr_t ae_id_lookup (ae_op_id_t id, int * resource_id);
+----
-A resource API uses the same types and conventions as the underlying resource; we don't try to hide any of that. It just handles how to post and complete blocking operations.
+The function returns the private data, and stores the resource ID of the
+resource creating the operation in `*resource_id`.
-Aesop is re-entrant and uses threads.
+Capturing the caller's Execution Context
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-A resource can use a number of progress models:
+Within a function created by the `ae_define_post` macro, it is possible to use
+the `ae_op_fill` macro to capture the aesop execution context. This macro
+takes the internal aesop parameters (similar to the function stack) and stores
+them in a `struct ae_op`.
-if the resource has its own threads or progress engine, then it can trigger callbacks that drive the next aesop execution state
-if the resource is passive, it can request polling from aesop and aesop will drive it with explicit poll calls
-Polling resources can "busy poll" or just indicate specific times when they would like to be polled
-Creating a new resource
-Look at resources/timer/timer.c as an example.
+The `ae_op` structure contains two members that can be read by the resource.
-The most important file that a resource will use to help define its interface is resource.h, which can be found in the top level directory.
+[horizontal]
-The ae_resource struct defines the interface to each resource. It includes the following pointers:
+*user_ptr*:: An internal data member used by aesop.
-resource_name
-test()
-poll_context()
-cancel()
-register_context()
-unregister_context()
-config_array()
-resource_name is the only mandatory field. The others are optional depending on what functionality your resource provides.
+*callback*:: This parameter is a pointer to a function `void (*callback) (void
+*, T)`, where `T` is the return type of the aesop blocking function being
+completed. The first parameter (`void *`) needs to be set to the `user_ptr`
+when calling the function.
-SSM as an example
-In ssm, the user calls a function called ssm_wait() which will trigger callbacks. ssm_wait() takes a timeout argument to tell it how long to wait. The callbacks are executed in serial in the context of the wait() call. wait() does not spawn threads. The callback functions can do pretty much anything; they can even post new SSM operations.
+Completing a blocking function can then be done by calling the callback
+provided by the stored in the `ae_op` structure (and filled in by the
+`ae_op_fill` macro).
-SSM init function returns a handle. A use case for calling init twice and getting two handles would be if you wanted to use two transports simultaneously.
+WARNING: The `ae_op` structure can be modified by the callback in certain
+cases. It is required to copy the `callback` and `user_ptr` values onto the
+stack *before* completing the blocking function.
-SSM makes progress on communication autonomously, even if you don't call wait(). So wait() does not drive communication progress, it only lets you find completion events.
+For example (from ResourceBuilder):
-If the ssm_wait() function is sleeping in one thread, while another thread registers a callback and does a put/get, then the wait _will_ pick up the new completion event. You don't have to restart the wait() call. This simplifies the resource greatly.
+----
+void (*callback) (void *, int) = op.callback; /* op = struct ae_op */
+void * user_ptr = op.user_ptr;
+callback (user_ptr, returncode);
+----
+
+Returning data from a blocking call
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+Since the return code of a blocking call defined using `ae_define_post` is
+always an integer, used by aesop internally, controlling the return value from
+a blocking call cannot be done using a simple `return`.
+
+If the call completed immediately (and the blocking function returns
+`AE_IMMEDIATE_COMPLETION`) or failed to start (returning an aesop error), the
+return code for the blocking call can be set through the `__ae_retval`
+pointer. This pointer is set by `ae_define_post`, and is only accessible
+within its scope. The type of the pointer is the same as the type listed as
+the first parameter to the `ae_define_post` function.
-Kevin's example of an SSM resource
+If the function was completed at a later point (by calling the aesop provided
+callback), the return code of the blocking function can be passed in as the
+second parameter to the callback (the first parameter is used by aesop
+internally).
-General plan: Kevin will provide a basic, possibly poor performing ssm resource, UAB team will own it from there to test performance and tune it, change threading, etc. to match best practice for SSM performance.
-Issues: we have to decide (soon, not necessarily today) where to host resources. Should ssm be part of the aesop repo, or should there be separate repos so that not every aesop user gets ssm, etc.
+// Changing the heading below to the other breaks the build???!?!?
+
+[[ref-register-resource]]
+==== Registering a New Resource
+
+The header `resource.h` defines the following structure:
+
+----
+struct ae_resource
+{
+ const char *resource_name;
+ int (*test)(ae_op_id_t id, int ms_timeout);
+ int (*poll_context)(ae_context_t context);
+ int (*cancel)(ae_context_t ctx, ae_op_id_t id);
+ int (*register_context)(ae_context_t context);
+ int (*unregister_context)(ae_context_t context);
+ struct ae_resource_config* config_array; /* terminated by entry with NULL name */
+};
+----
-Code walkthrough
+* The `resource_name` member is the only mandatory field in this structure.
-There are some minor differences between "in tree" version of aesop within the triton repo, and the "stand alone" version of aesop that we are working with. Kevin's example is in tree, and will need some minor porting to go along with aesop.
+* The `test` function, if implemented, can be used by aesop to wait for the
+ completion of a specific operation, specified by the `id` parameter.
+ Generally, it is not necessary to implement this function; the `test`
+ function pointer can be set to `NULL`.
-Error codes: functions that aesop actually uses directly (init and finalize are good examples) you must use pre-defined error codes. For functions that are specific to your resource (like put() and get()) you can do absolutely anything that you want.
+* The `cancel` function is an asynchronous request to cancel the specified
+ operation. The resource is still responsible for completing the blocking
+ function by calling the operation's callback from the `ae_op` structure.
-There is a call to register the initialization and finalization routines (triton_init_register()).
+* The `poll` and `config_array` members will be discussed below.
-The initialization function: register the resource with aesop, specify the ae_resource struct that fills in function pointers for various functionality. Then you create a default context.
-Right now the transport and address are hard coded (using tcp on localhost).
+[NOTE]
+====
+The _context_ concept is deprecated and will be removed in a future version.
+The `register_context` and `unregister_context` functions should be set to
+`NULL`. The `poll_context` function can safely ignore the `context` parameter.
+====
-ae_define_post(... triton_ssm_put ...)
-The ae_define_post lets you specify a blocking function and its arguments. It (behind the scenes) tacks on extra arguments that are needed by aesop.
+The new resource should be registered with aesop using the `ae_resource_register`
+function. The internal ID for the resource (used in creating an operation ID)
+is returned in `*newid`.
-Blocking functions like this can support immediate completion. SSM does not do this yet, but it is something we can discuss later as an optimization. The idea is to avoid context switch to another thread if you post an operation that can be finished in place.
+----
+int ae_resource_register (struct ae_resource *resource, int *newid);
+void ae_resource_unregister (int resource_id);
+----
-The opcache is an aesop thing that lets you allocate a struct to represent an in-flight operation (an "op"). It has a user-definable field that you can use to tack on information specific to your resource.
+[TIP]
+====
+The `ResourceBuilder` and `Thread` resources provide good examples on how to
+construct a simple resource. Their code can be found in the `resources`
+directory in the aesop distribution.
+====
-The following macro populates the op structure with fields to tell aesop what to do when the operation completes:
+[[ref-resource-configuration]]
+Resource Configuration
+~~~~~~~~~~~~~~~~~~~~~~
-ae_op_fill_with_params()
-General comment: this example needs one line comments explaining what's going on, and point out which things are optional.
+When a resource is registered with the aesop system using the
+`ae_resource_register()` function, the resource author has the option of
+specifying configuration parameters for that resource using the
+`config_array` field of the required `ae_resource` struct. The
+`config_array` field is a `NULL` terminated array of structs of type `struct
+ae_resource_config`, which is defined as follows:
-General comment: it might be a good idea for ANL to just implement some basic functions and then hand off to UAB to fill in remainder, would be a good exercise for everyone.
+----
+struct ae_resource_config
+{
+ const char* name;
+ const char* default_value;
+ const char* description;
+ int (*updater)(const char* key, const char* value);
+};
+----
-This function can be used to track operations (put them in whatever queues you would like as a resource imlementor):
+The `name` field is the name of the configuration parameter (which
+corresponds to the `key` argument of the `aesop_set_config()` function).
+The `default_value` field specifies the starting value of the configuration
+parameter. The `description` is a free form text field, though standard
+practice is to limit this to a single line of text describing the
+configuration parameter. The `updater` field is a pointer to a function
+that will be invoked when the configuration parameter is updated. The
+resource must provide this function and handle configuration updates in a
+manner appropriate for the resource. For example, this function may
+perform an `sscanf()` to read an integer value, and then check its values
+against the legal ranges for the parameter before accepting it. The
+resource should return 0 on success or `AE_ERR_INVALID` if the configuration
+value is not valid.
+
+The `config_array` field can be set to `NULL` at resource registration time
+if the resource does not support any configuration parameters.
-ae_ops_enqueue()
-The caller of ae_ops_enqueue() is responsible for appropriate locking when modifying or moving op structures around. Until the resource completes an operation (and hands off control to aesop) it is the resource's responsibility to handle them until then.
+[NOTE]
+====
+As of February 2012 this functionality is only avaialable to resources, but
+in future work we will extend this concept to allow arbitrary Aesop
+components to register configuration parameters in a similar manner.
+====
-The following function handles polling:
-triton_ssm_poll()
-Right now the resource busy spins and expects aesop to call the poll function constantly. We know that this is not a final implementation. In the longer term we want this resource to have a thread that drives the wait() function. Once that is in place then the poll function is no longer needed.
+[[ref-progress]]
+Resource Progress
+~~~~~~~~~~~~~~~~~
-General issue: we need to discuss whether to keep wrappers for things like triton_mutex_lock(). If we do want to keep wrappers, we need to decide whether each component does its own wrappers or we all agree on a centralized implementation/wrapper across the project.
+A resource can use a number of progress models:
-Future work (not enough time in this session)
-Need to address semantics of ssm in relation to anl/triton work, independent of the resource implementation. Let's pick back up on that this afternoon after completing the agenda.
+First, if the underlying API exported by the resource has its own threads or
+progress engine, then it can trigger callbacks that drive the next aesop
+execution state. The POSIX AIO functionality is an example of this.
-UAB can also send visitors to ANL easily if we need more interaction later.
+If the underlying functionality is passive, the resource can request polling
+from aesop and aesop will drive it by explicit calls to the resource polling
+function specified at registration time. Polling resources can "busy poll" or
+just indicate specific times when they would like to be polled.
-////
+A resource is responsible for indicating it has work to do, by calling the
+`ae_resource_request_poll` function. After calling this function, aesop will
+schedule a call to the resource's `poll_context` function.
+The `ae_resource_request_poll` function is thread-safe and can safely be
+called from within a signal handler.
+A combination of these models is also possible. For example, a signal handler
+or callback can record the completion of a resource function, and request a
+poll. When the resource's polling function is subsequently called by aesop,
+the resource can complete the action by calling the appropriate aesop
+callback.
When to write a resource
~~~~~~~~~~~~~~~~~~~~~~~~
@@ -216,3 +494,4 @@ requires less code, these modules already take care of dealing with low level
details such as generating operation id's and cancellation.
+
diff --git a/doc/aug_standard.txt b/doc/aug_standard.txt
index 3ee81ad..1a73dfd 100644
--- a/doc/aug_standard.txt
+++ b/doc/aug_standard.txt
@@ -6,14 +6,15 @@ The Aesop Standard Library
Aesop System Interfaces and Tools
----------------------------------
-
Configuration Interface
~~~~~~~~~~~~~~~~~~~~~~~
Aesop provides a unified interface to advertise and set configuration
-parameters across all registered resources. This section describes both
-the Aesop user interface to this functionality as well as the Aesop library
-author's interface to this functionality.
+parameters across code modules.
+
+This section describes the Aesop user interface to this functionality; The low
+level C interface used for system components is described in
+<<ref-resource-configuration>>.
Querying and setting configuration parameters
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
@@ -35,46 +36,6 @@ There is currently no corresponding function to query a configuration
parameter or list available keys (as of February 2012), although this will be
added in future versions of Aesop.
-Registering new configuration parameters
-^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
-
-When a resource is registered with the Triton library using the
-`ae_resource_register()` function, the resource author has the option of
-specifying configuration parameters for that resource using the
-`config_array` field of the required `ae_resource` struct. The
-`config_array` field is a `NULL` terminated array of structs of type `struct
-ae_resource_config`, which is defined as follows:
-
-----
-struct ae_resource_config
-{
- const char* name;
- const char* default_value;
- const char* description;
- int (*updater)(const char* key, const char* value);
-};
-----
-
-The `name` field is the name of the configuration parameter (which
-corresponds to the `key` argument of the `aesop_set_config()` function).
-The `default_value` field specifies the starting value of the configuration
-parameter. The `description` is a free form text field, though standard
-practice is to limit this to a single line of text describing the
-configuration parameter. The `updater` field is a pointer to a function
-that will be invoked when the configuration parameter is updated. The
-resource must provide this function and handle configuration updates in a
-manner appropriate for the resource. For example, this function may
-perform an `sscanf()` to read an integer value, and then check its values
-against the legal ranges for the parameter before accepting it. The
-resource should return 0 on success or `AE_ERR_INVALID` if the configuration
-value is not valid.
-
-The `config_array` field can be set to `NULL` at resource registration time
-if the resource does not support any configuration parameters.
-
-As of February 2012 this functionality is only avaialable to resources, but
-in future work we will extend this concept to allow arbitrary Aesop
-components to register configuration parameters in a similar manner.
Debugging
~~~~~~~~~~
@@ -111,8 +72,10 @@ The `value` parameter can be set to `1` or `0`, in order to respectively
enable or disable tracing.
+//=====================================================================
Standard Aesop modules
-----------------------
+//=====================================================================
The following modules are bundled with the aesop distribution.
@@ -309,6 +272,7 @@ thread
WARNING: The thread module relies on internal details of the current aesop
implementation. Its use and interface might change in future versions.
+Header: `aethread.hae`
The following functions are used to intialize and shut down the thread module.
hooks/post-receive
--
aesop Repository
1
0
1
0
aesop Repository branch, master, updated. 66ac30aaaf1f4b1174b29c5bb89a19cc3491784d
by noreply@mcs.anl.gov 27 Feb '12
by noreply@mcs.anl.gov 27 Feb '12
27 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via 66ac30aaaf1f4b1174b29c5bb89a19cc3491784d (commit)
from 581027046f89af24e644f93489187a7f5c8a4905 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 66ac30aaaf1f4b1174b29c5bb89a19cc3491784d
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Mon Feb 27 17:06:47 2012 -0500
brief documentation for the aesop config interface
-----------------------------------------------------------------------
Summary of changes:
doc/aug_standard.txt | 64 ++++++++++++++++++++++++++++++++++++++++++++++++++
1 files changed, 64 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/doc/aug_standard.txt b/doc/aug_standard.txt
index 0817c8f..3ee81ad 100644
--- a/doc/aug_standard.txt
+++ b/doc/aug_standard.txt
@@ -10,7 +10,71 @@ Aesop System Interfaces and Tools
Configuration Interface
~~~~~~~~~~~~~~~~~~~~~~~
+Aesop provides a unified interface to advertise and set configuration
+parameters across all registered resources. This section describes both
+the Aesop user interface to this functionality as well as the Aesop library
+author's interface to this functionality.
+Querying and setting configuration parameters
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+The following function can be used to set arbitrary configuration
+parameters:
+
+----
+int aesop_set_config(const char* key, const char* value);
+----
+
+The `key` argument is the name of the configuration parameter to be set,
+while the `value` argument is the new value for that configuration
+parameter. All values must be provided as strings. The function returns 0
+on success, `AE_ERR_NOT_FOUND` if the key is not known to Aesop, or
+`AE_ERR_INVAL` if the key is valid but the value is not.
+
+There is currently no corresponding function to query a configuration
+parameter or list available keys (as of February 2012), although this will be
+added in future versions of Aesop.
+
+Registering new configuration parameters
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+When a resource is registered with the Triton library using the
+`ae_resource_register()` function, the resource author has the option of
+specifying configuration parameters for that resource using the
+`config_array` field of the required `ae_resource` struct. The
+`config_array` field is a `NULL` terminated array of structs of type `struct
+ae_resource_config`, which is defined as follows:
+
+----
+struct ae_resource_config
+{
+ const char* name;
+ const char* default_value;
+ const char* description;
+ int (*updater)(const char* key, const char* value);
+};
+----
+
+The `name` field is the name of the configuration parameter (which
+corresponds to the `key` argument of the `aesop_set_config()` function).
+The `default_value` field specifies the starting value of the configuration
+parameter. The `description` is a free form text field, though standard
+practice is to limit this to a single line of text describing the
+configuration parameter. The `updater` field is a pointer to a function
+that will be invoked when the configuration parameter is updated. The
+resource must provide this function and handle configuration updates in a
+manner appropriate for the resource. For example, this function may
+perform an `sscanf()` to read an integer value, and then check its values
+against the legal ranges for the parameter before accepting it. The
+resource should return 0 on success or `AE_ERR_INVALID` if the configuration
+value is not valid.
+
+The `config_array` field can be set to `NULL` at resource registration time
+if the resource does not support any configuration parameters.
+
+As of February 2012 this functionality is only avaialable to resources, but
+in future work we will extend this concept to allow arbitrary Aesop
+components to register configuration parameters in a similar manner.
Debugging
~~~~~~~~~~
hooks/post-receive
--
aesop Repository
1
0
aesop Repository branch, master, updated. 581027046f89af24e644f93489187a7f5c8a4905
by noreply@mcs.anl.gov 27 Feb '12
by noreply@mcs.anl.gov 27 Feb '12
27 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via 581027046f89af24e644f93489187a7f5c8a4905 (commit)
from 32c4c7bf45f3be6542fdd0e99b6aabcfa85beffe (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 581027046f89af24e644f93489187a7f5c8a4905
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Mon Feb 27 16:41:34 2012 -0500
brief notes on how to use the aecc translator
-----------------------------------------------------------------------
Summary of changes:
doc/aug_introduction.txt | 20 ++++++++++++++++++++
1 files changed, 20 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/doc/aug_introduction.txt b/doc/aug_introduction.txt
index de1e557..a80c83e 100644
--- a/doc/aug_introduction.txt
+++ b/doc/aug_introduction.txt
@@ -325,6 +325,26 @@ additional ae.i and ae.s files to be created during the compilation of aesop cod
These files contain the C translation of the aesop source code.
=====
+Invoking the Aesop translator (aecc)
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The following example shows the command line used to translate a .ae file
+into an object file. The resulting object file can be linked using `ld` or
+`gcc` just as you would link an object produced by the standard gcc compiler.
+
+----
+aecc -o file.o file.ae -c gcc -- -g
+----
+
+The `-c` argument specifies the C compiler to use, which must be a version
+of gcc. The `--` argument is a delimiter indicating that all following
+arguments (`-g` in this case) will be passed directly to the C compiler.
+
+As of February 2012 there are some known bugs in using the aecc translator
+outside of the Aesop repository. More details can be found in trac
+tickets #21 and #7. Until these are resolved we recommend using the test
+programs in the Aesop tree as an example of how to best set the `-I` include
+path arguments and `-L` library path arguments when compiling external code.
// @TODO Explain problem using gdb
hooks/post-receive
--
aesop Repository
1
0
I believed we were asked to provide feedback on the new SSM API. Here are our comments and questions.
kevin
Argonne SSM API Feedback
1. [required] ssm_put/ssm_get missing length parameters.
- The local buffer is registered in its entirety. Example: local_buf, 30MB, the ssm_put/get should specify the local offset within this buffer and the total amount of data to send, for example, offset 10 MB. I'm not sure if we need to specify the length for both the local and remote buffer or we can just specify the length once.
2. [required] ssm_main_addr.h -- there should be a ssm_addr_compare function.
- A function is need to compare two ssm_addr_id pointers to see if they reference the same address.
3. [comment] offset parameters -- these should be off_t
- i think off_t would be more appropriate instead of size_t
4. [comment] Inconsistent API
- The ssm_cancel/ssm_wait function requires specifying the 'ssm id' but the put/get/msg don't require it. (referenced from match entry) Since cancel requires the id, just move this parameter to all operations.
5. [comment] ssm_put/get/msg APIs are not intuitive.
- The current ssm_put and ssm_get APIs take a match entry. This is a little non-intuitive because this data isn't defining an actual match entry. In this case the match entry just serves to aggregate the ssm id, local buffer description, destination match bits, and callback. It might be better served to just make these parameters to put and get or define a new local buffer data structure specifically for put/get/msg operations.
6. [comment] ssm_me_add only allows adding one buffer per call.
- Requires looping over ssm_me_add to multiple buffers. This call or another variant could allow adding an array of buffers.
ssm_me_add (ssm_me me, void ** buf, const size_t * sizes)
(We assume SSM will always take the next free buffer off the list)
Questions
A. Can we setup concurrent put/gets using the same match entry?
- is this considered safe?
me = ssm_me_new()
ssm_link(me)
ssm_put(me, svr1)
ssm_put(me, svr2)
ssm_put(me, svr3)
B. If we are using never unlink, during the completion callback, do we add the buffer back with ssm_me_add() or do we use a different mechanism?
2
1
aesop Repository branch, master, updated. 32c4c7bf45f3be6542fdd0e99b6aabcfa85beffe
by noreply@mcs.anl.gov 27 Feb '12
by noreply@mcs.anl.gov 27 Feb '12
27 Feb '12
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "aesop Repository".
The branch, master has been updated
via 32c4c7bf45f3be6542fdd0e99b6aabcfa85beffe (commit)
via ed7959e5255a60a1a59c49c9d18605b0ee7e45b3 (commit)
via e668851e4054586670fb44cc19c3a205b66af19f (commit)
from 497a206ef3c8944344fbcbb4079d71f4bb9d218d (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 32c4c7bf45f3be6542fdd0e99b6aabcfa85beffe
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Mon Feb 27 10:32:40 2012 -0500
spell check
commit ed7959e5255a60a1a59c49c9d18605b0ee7e45b3
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Mon Feb 27 10:26:28 2012 -0500
misc editing of remainder of aesop perf doc
commit e668851e4054586670fb44cc19c3a205b66af19f
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Mon Feb 27 10:07:53 2012 -0500
shuffle around some sections of text
-----------------------------------------------------------------------
Summary of changes:
doc/aesop-performance.txt | 365 +++++++++++++++++++++++----------------------
1 files changed, 187 insertions(+), 178 deletions(-)
Diff of changes:
diff --git a/doc/aesop-performance.txt b/doc/aesop-performance.txt
index e182ba7..942fd2e 100644
--- a/doc/aesop-performance.txt
+++ b/doc/aesop-performance.txt
@@ -9,23 +9,23 @@ performance.
Aesop is a programming language and programming model designed
to implement distributed system software with high development productivity and run
time efficiency.
-Further details about the languange and development environment can be found
+Further details about the language and development environment can be found
in the Aesop User's Guide.
The remainder of this document is organized as follows.
-<<sec-case-study>> and <<runtime-perf>> describe a network service
-case study and use it to evaluate the performance of Aesop
-relative to more traditonal server architectures.
-<<sec-memory>> provides a more detailed analysis of memory efficiency in
-Aesop. <<sec-productivity>> evaluates the productivity of the Aesop
-language using the original case study as an example.
+<<sec-case-study>> describe a network service
+case study which is then used to evaluate the performance of Aesop
+relative to more traditional server architectures in terms of performance
+(<<runtime-perf>>), memory usage (<<sec-memory>>), and productivity
+(<<sec-productivity>>). <<sec-overhead-analyis>> provides a more detailed
+breakdown of specific sources of Aesop overhead, while
<<sec-compile>> concludes by discussing Aesop compile-time code translation
performance.
[[sec-case-study]]
== Case study description
-We will use a simple network serivce case study for quantitative
+We will use a simple network service case study for quantitative
evaluation of the Aesop programming language. The case study is a TCP
server that can write or read data from local files. It is expected to
process requests from multiple clients simultaneously. The description
@@ -40,10 +40,10 @@ client test harness. The client is a basic C program that uses TCP sockets to s
the server. It use MPI to coordinate processes and generate a highly
concurrent workload.
-The client will execute in a loop generating a specifed number of
+The client will execute in a loop generating a specified number of
operations to the server. The general flow is that each client process sends a request
to the server that contains an optional payload. The client then waits for
-the server to send an acknowledgement that also contain an optional payload.
+the server to send an acknowledgment that also contain an optional payload.
==== Request Types
@@ -53,13 +53,13 @@ The client supports the following request types.
The client sends a request with a file name and a size. The server will then
open the file, read the contents up to the size specified. The server returns
-the data with the acknowledgement of the operation.
+the data with the acknowledgment of the operation.
===== Write
The client sends a request with a file name, size and payload. The server will
then create the file and write the payload. The server then sends an
-acknowledgement to client.
+acknowledgment to client.
===== Read-Null
@@ -67,21 +67,21 @@ This request is identical to the *Read* request, except that the server
sends uninitialized data rather than performing any file I/O.
The client sends a request with a size to the server. The server then allocates
a buffer for the response based on the size the client requested. The server
-then sends an acknowledgement with this buffer as the payload.
+then sends an acknowledgment with this buffer as the payload.
===== Write-Null
This request is identical to the *Write* request, except that the server
discards incoming data rather than performing any file I/O.
-The client sends a request with a size and a payload. The server recieves
+The client sends a request with a size and a payload. The server receives
the request but then simply discards the payload. The server sends an
-acknowledgement back to the client.
+acknowledgment back to the client.
[float]
==== Implementation
The client test harness provides command line parameters to control the
-request type, the number of requets, and the size of the request. Each
+request type, the number of requests, and the size of the request. Each
process barriers until they are all ready to
connect to the server. Clients then exit the barrier, connect to the server,
and begin sending requests in a loop. Once each process completes its
@@ -89,10 +89,10 @@ requests, it waits
at another barrier and then reports statistics about the run. Each process
records the total amount of time taken to execute its workload (beginning
before the initial connection and ending after receipt of the last
-acknowledgement). The time taken by the slowest process is reported as the
+acknowledgment). The time taken by the slowest process is reported as the
aggregate run time. Each process also records the time needed to service each
individual request (from before the request is sent until after the
-acknowledgement is received) in order to calculate statistics about
+acknowledgment is received) in order to calculate statistics about
individual request latencies.
=== Server Design
@@ -195,7 +195,7 @@ is handled completely from within that thread. When the request is complete
the thread is destroyed. Blocking socket operations and standard file read
and write functions are used in this implementation.
-===== Thead-pool
+===== Thread-pool
The thread-pool server uses an event loop to watch all sockets for activity.
When a new request is available, the event loop puts the request on a queue and
wakes up a thread from the thread pool. The request is handled completely
@@ -233,10 +233,10 @@ that disk I/O is _not_ involved in those cases.
All experiments were executed on the Fusion cluster managed by
the Argonne Laboratory Computing Resource Center
-(LCRF). Fusion is a IBM iDataPlex dx360 M2 system. It features 320 compute
-nodes which each consist of two Intel Nehalem 2.6 GH Xeon processors and 36 GB
+(LCRC). Fusion is a IBM iDataPlex dx360 M2 system. It features 320 compute
+nodes which each consist of two Intel Nehalem 2.6 GHz Xeon processors and 36 GB
of RAM. The compute nodes have hyper threading disabled. The cluster has
-an Infiniband QDR interconnect. Each compute node also a single SATA 7200 RPM
+an InfiniBand QDR interconnect. Each compute node also a single SATA 7200 RPM
hard disk for local scratch storage.
==== Experiment Details
@@ -279,7 +279,7 @@ performed with TCP/IP sockets.
The read test had clients each issue 16 requests asking for 4 KiB from
disk. Each client specifies a unique file to be read on on each request. All
-clients specifiy unique files. The files are first generated by a script
+clients specify unique files. The files are first generated by a script
that runs before the read test starts. The script generates files for every
client in the local storage of the server.
@@ -292,7 +292,7 @@ mpirun -np <procs> echo-client --ip <ip> --port 9999 --path <path> --num-request
The write test had clients each issue 16 requests sending 4 KiB of data to
be written to disk. Each client specifies a unique file name for each request
-and all clients specifiy unique files from each other. The directory
+and all clients specify unique files from each other. The directory
containing all the files is deleted between each test iteration.
.Execution Parameters
@@ -349,7 +349,7 @@ perform as well as the other servers for small workloads (taking 3.2
seconds at the smallest scale, verses 1.9 seconds for the thread-per-op
server). However, Aesop is the fastest server at the largest scale
(taking 119.8 seconds verses 130.8 seconds for the nearest competitors
-in thread-per-client and threhad-per-client-nb).
+in thread-per-client and thread-per-client-nb).
.Runtime Performance for Read Test
[[fig-readhist]]
@@ -365,14 +365,14 @@ cases, ultimately running the largest scale test in 77.2 seconds.
The small scale results for Aesop may indicate that additional tuning
is needed to improve latency for small test runs. The issue is likely
isolated to the write path of the file I/O resource in the Aesop standard
-library, as we see assymetric results in the read and write tests for Aesop
+library, as we see asymmetric results in the read and write tests for Aesop
in terms of its relative performance.
==== Network I/O
The write-null and read-null experiments were conducted in the same manner
as the write and read tests. The difference in this case is that no disk
-access was perfomed. In the write case, incoming data was discarded by the
+access was performed. In the write case, incoming data was discarded by the
server. In the read case, the server transmitted uninitialized data.
.Runtime Performance for Write-Null Test
@@ -399,7 +399,7 @@ thread-per-client server to the point that it is practically equivalent to
the Aesop server at scale.
Another notable observation in these graphs is that the Aesop server is
-competative at small scale, and in fact is the fastest implementation in the
+competitive at small scale, and in fact is the fastest implementation in the
16 client process read-null test and nearly the fastest in the 16 client
process write-null test. This supports the observation from the previous
section that poor Aesop performance at small scale is likely a tuning flaw
@@ -468,9 +468,9 @@ metrics are graphed using a box and whiskers plot. The box represents the
first and third quartiles and the whiskers are the minimum and maximum
values. The following results are for the 1024 client size selected from
the same iteration as the maximum runtime graphs. <<fig-readlat>> and
-<<fig-writelat>> show the aesop offers vary comparitve latency performance
-as the other configurations and only noteably thread-per-op and event are
-signficantly worse. An interesting observation is that the overall fairness
+<<fig-writelat>> show the aesop offers vary comparative latency performance
+as the other configurations and only notably thread-per-op and event are
+significantly worse. An interesting observation is that the overall fairness
across clients shown in the previous section does not appear to correlate in
any way with individual request latency.
@@ -492,20 +492,20 @@ implementations.
The runtime analysis shows that aesop is competitive with the most
optimal alternate server implementations at the 1024 client size.
In the _read_ test, aesop is 4.5% slower that the fastest implementation,
-threadpool. Aesop was the fastest implementation in the _write_ test.
+thread-pool. Aesop was the fastest implementation in the _write_ test.
The thread-per-client implementation was fastest in the _read null_ and
_write null_ tests. Here aesop is significantly slower at around 25-30%,
however, as we showed above, the difference in performance was due to the
native performance difference in synchronous and asynchronous sockets.
If we look at the performance comparison to the thread-per-client-nb, aesop
was only 4.7% slower in the _read null_ test and was faster in the _write null_
-test. If aesop were using a native asynchronous transport such as Infiniband verbs
+test. If aesop were using a native asynchronous transport such as InfiniBand verbs
or SSM, we believe that it would perform as well as a thread-per-client implementation.
Another interesting factor was that there was no one fastest server
implementation. Each workload presents a different challenge and building
a single tuned server is difficult. A key advantage of aesop is the ability
-to change the underyling implementation/turning for resources or the aesop
+to change the underlying implementation/turning for resources or the aesop
runtime without changing the aesop server source. During our investigation
we experimented with different underlying thread models for the network and
file resources of aesop. The aesop resources can be configured at runtime
@@ -514,21 +514,143 @@ that building a specific server implementation that is optimal for a generic
workload, aesop becomes powerful because these types of changes can be made
without ever changing the server source.
-=== Analysis
+[[sec-memory]]
+== Runtime Memory Efficiency
+
+Another aspect of the overall performance is the memory efficiency of each
+server implementation. In this section we compare aesop to the other server implementations
+as we did for the runtime performance.
+
+=== Experiment
+
+The memory utilization of each server implementation was captured during
+the performance experiments detailed in <<runtime-perf>>. We recorded
+the VmHWM stat from the server when the client test was completed. The
+VmHWM stat is a Linux-specific metric that represents the peak resident
+set size (RSS) of an executable, where RSS corresponds to the amount
+of paged-in memory used by the executable.
+
+=== Evaluation
+
+The following graphs show the memory usage in KiB in log scale.
+
+.Memory Usage Write Test
+[[fig-writemem]]
+image::fig/write-mem.png[]
+
+.Memory Usage Read Test
+[[fig-readmem]]
+image::fig/read-mem.png[]
+
+In <<fig-readmem>> and <<fig-writemem>> we see that thread-pool limits the
+memory usage as the client work load increases because the thread-pool
+by design limits the number of requests that can be in progress at once. The
+other server implementations scale as the number of clients increase.
+
+.Memory Usage Write-Null Test
+[[fig-writenullmem]]
+image::fig/write-null-mem.png[]
+
+.Memory Usage Read-Null Test
+[[fig-readnullmem]]
+image::fig/read-null-mem.png[]
+
+<<fig-readnullmem>> and <<fig-writenullmem>> show a similar result as the
+disk I/O tests. In this case the event server also produces favorable
+results (in addition to the thread pool server) because it is the only
+implementation which does not spawn any additional threads to handle
+increasing numbers of network connections.
+
+Although the Aesop server cannot match the thread-pool server in terms of
+memory usage, it does compare favorably to the thread-per-client and
+thread-per op servers. In the read test, for example, Aesop consumes
+roughly 10 MiB of paged in memory verses almost 18 MiB of paged in memory
+for the thread-per-client server at scale.
+
+Note that the thread-per-client and thread-per-op models consume virtual
+memory at a much larger rate due to the number of thread stacks allocated.
+We chose not to evaluate this metric, however, as the resident memory seems
+to be a more relevant metric in practice.
+
+[[sec-productivity]]
+== Productivity
+
+The core design element of aesop is to make programming of a concurrent
+server easier, so the trade off for memory and runtime performance should
+be worth it. To evaluate this we examine the code complexity of each of
+server implementations.
+
+.Implementation complexity analysis.
+[[table-complex]]
+[cols="3,1,1,1", options="header"]
+|============================
+| Server Implementation | CC | Mod. CC | SLOC
+| aesop | 16 | 11 | 179
+| thread-per-client | 17 | 12 | 182
+| thread-per-client-nb | 17 | 12 | 184
+| thread-per-op | 22 | 17 | 249
+| thread-pool | 32 | 26 | 313
+| event | 28 | 23 | 341
+|============================
+
+<<table-complex>> compares the code complexity of each server
+implementation using McCabe Cyclomatic Complexity (CC) <<McCabe>>,
+Modified McCabe Cyclomatic Complexity (Mod. CC), and Source Lines of Code
+(SLOC). The CC and Mod. CC metrics were measured using the pmccabe tool,
+version 2.6, created by Paul Bame <<Bame>>, while the SLOC metrics were
+measured using the sloccount tool, version 2.26,
+created by David A. Wheeler <<Wheeler>>.
+
+To simplify the comparison, error handling was excluded for all servers
+except for assertions on expected return codes. The protocol definition
+(ie, request and acknowledgment structs) as well as helper functions to
+loop over send and receive were not counted in any of the implementations,
+as these were similar in all four. Also note that the Aesop version
+does not include the Aesop standard library, which
+provides a binding between Aesop and the standard POSIX socket API as part
+of its default functionality. The focus
+of this comparison is on the core logic defining the server implementation.
+
+The Aesop and thread-per-client servers are very similar in terms of
+complexity. The slight increase in complexity for the thread-per-client
+server results from
+the additional function calls needed to create and join threads. The
+non-blocking version of the thread-per-client server (thread-per-client-nb)
+uses two additional lines of code to place each socket into non-blocking
+mode. The remaining code logic needed to manage the non-blocking socket
+calls is implemented in helper functions which are not included in the
+analysis.
+
+The thread-pool and event models are both much more complex than the
+thread-per-client or aesop model. In the case of the thread-pool server,
+this additional complexity arises from not only the queuing and thread
+management logic, but also the event loop which is necessary to detect
+incoming requests and dispatch them to the queue. The event server
+complexity arises from the necessity of dividing servicing routines
+into multiple sub-functions and manually tracking state between
+those functions. An additional complexity of the event
+model which is not captured by these metrics is the fact that control flow
+is not preserved across the processing of a given request. For example,
+servicing a write operation requires 5 disconnected event handlers.
+Although the event model appears less complex than the thread-pool model according to CC and
+Mod. CC, qualitatively it is significantly more challenging to develop.
+
+[[sec-overhead-analysis]]
+== Sources of Aesop overhead
As we've seen aesop compares favorably to other concurrency models but loses
some performance compared to best case hand tuned version. Here we examine
the overheads associated with aesop. The first item to examine is the
cost associated with an aesop blocking call.
-==== Performance Implications of Blocking Calls
+=== Performance Implications of Blocking Calls
While this isn't immediately visible from looking at the aesop source code,
blocking calls, when compared to a plain C function call, have extra overhead
due to the way they are transformed by the aesop compiler. The following
section highlights the sources of this overhead.
-===== State Management
+==== State Management
Most of the overhead is caused by the need to preserve the state
of the blocking call while execution (temporarily) switches to another
@@ -539,7 +661,7 @@ variables. As allocating heap memory is much more time consuming than
allocating space on the stack, calling a blocking function is more expensive
than calling a regular function.
-===== Synchronization Overhead
+==== Synchronization Overhead
A second source of overhead originates from the multi-threaded nature of aesop
code. While aesop does not create any threads, many of the aesop resources
@@ -570,13 +692,13 @@ Currently, the aesop compiler uses a combination of atomic operations and
mutexes to maintain thread-safety. There is an ongoing effort to convert to
atomic operations where possible.
-===== Quantifying Blocking Call Overhead
+==== Quantifying Blocking Call Overhead
For this test, a regular C function and a aesop blocking function are
called in a loop. By timing the total time required to complete the loop,
an estimate of the time needed to execute the function call is obtained.
-The results were obtained on an intel i7 CPU running at 2.7GHz,
+The results were obtained on an Intel i7 CPU running at 2.7 GHz,
using gcc 4.5.3 (using +-O2+), glibc 2.13-r4 and kernel 3.2.5.
There are a number of different test configurations:
@@ -598,9 +720,9 @@ arguments is used. Test 2 uses the same function, but this time the
function is marked as `__blocking`.
Tests 3-5 were added to provide a better context for understanding the
-magnitude of the blocking call overhead. For test 3, the function from test 1
-was taken but in the function body a single call to +malloc+ and +free+ was
-added. Test 4 is the same as test 3, but also adds a call to lock and unlock
+magnitude of the blocking call overhead. For test 3, the normal C function from test 1
+was used as a basis, but a single call to +malloc+ and +free+ was
+added within the test function. Test 4 is the same as test 3, but also adds a call to lock and unlock
a mutex. Test 5 replaces the mutex by a single atomic operation
(compare-and-swap).
@@ -609,83 +731,33 @@ also executed using a +tcmalloc+, an alternative memory allocator library.
The results for these are shown in the +tcmalloc+ column.
[NOTE]
-The progam used to obtain these results is in the repository:
+The program used to obtain these results is in the repository:
+tests/blocking-overhead.ae+.
-[[sec-memory]]
-== Runtime Memory Efficiency
-
-Another aspect of the overall performance is the memory efficiency of each
-server implementation. We compare aesop to the other server implementations
-as we did for the runtime performance.
-
-=== Experiment
-
-During the experiment detailed in the <<runtime-perf>> section. We recorded
-the VmHWM stat from the server when the client test was completed. The VmHWM
-stat is recorded by Linux during an applications runtime and represents the
-peak resident set size (RSS). RSS represents the amount of paged-in memory.
-If we looked at the VmPeak which is the total required virtual memory, this
-would punish the models which use numerous threads.
-
-=== Evaluation
-
-The following graphs show the memory usage in KiB in log scale. In
-<<fig-readmem>> and <<fig-writemem>> we see that thread-pool limits the
-memory usage as the client work load increases because the thread-pool
-by design limits the number of requests that can be in progress at once. The
-other server implementations scale as the number of clients increase.
+From these results we see that an Aesop __blocking function is significantly
+slower than a basic C function. Test case 4 illustrates that almost all of
+this overhead is a result of the memory management and synchronization
+performed by Aesop in order to enable efficient concurrency.
-<<fig-readnullmem>> and <<fig-writenullmem>> show a similar result as the
-disk I/O tests. Aesop demonstrates that it is no worse then any of the thread
-models.
+It is also important to note that Aesop is a superset of the C language, and
+normal C functions (with the associated low overhead) can still be used in
+an Aesop program. The __blocking functions are most appropriate for
+functions that perform blocking device or resource operations, while
+standard C functions are better suited for computationally intense inner-loop routines.
-.Memory Usage Read Test
-[[fig-readmem]]
-image::fig/read-mem.png[]
+=== Memory usage implications of blocking calls
-.Memory Usage Write Test
-[[fig-writemem]]
-image::fig/write-mem.png[]
+The main memory overhead incurred by blocking functions originates
+from the need to protect the logical state of the function while
+temporarily switching to other functions. Stack variables, function
+arguments, and return values must all be moved to the heap by the
+aesop source translator. In a normal C program, those items consume
+stack space. In addition, aesop internally maintains a number of control
+structures. Pointers to these structures are passed as function arguments
+to the blocking function, and consequently consume stack space. Currently,
+aesop adds about 4 pointers and 2 integers to each blocking function call.
-.Memory Usage Read-Null Test
-[[fig-readnullmem]]
-image::fig/read-null-mem.png[]
-
-.Memory Usage Write-Null Test
-[[fig-writenullmem]]
-image::fig/write-null-mem.png[]
-
-=== Analysis
-
-Aesop does not exhibit worse scaling in terms of memory usage than any of
-the other thread-per models but it does add a cost in memory overhead for
-each blocking call.
-
-=== Function arguments and stack variables of blocking calls
-
-Aesop, in order to implement the additional functionality provided by blocking
-calls, rewrites blocking calls when translating the aesop code to C code.
-This translation introduces a certain amount of overhead, both in memory usage
-and execution performance.
-
-The main memory overhead incurred by blocking functions originates from the
-need to protect the logical state of the function while temporarily switching
-to other functions.
-
-For example, stack variables are moved to the heap. As long as the blocking
-function does not complete, the memory for these variables is not released.
-Arguments to the function need to be relocated to the heap as well,
-and so does the type returned from the function (if not void).
-
-In a normal C program, the items listed above consume stack space. In blocking
-functions, these consume heap space instead. In addition, aesop internally
-maintains a number of control structures. Pointers to these structures are
-passed as function arguments to the blocking function, and consequently
-consume stack space. Currently, aesop adds about 4 pointers and 2 integers to
-each blocking function call.
-
-=== Lonely pbranches
+==== Lonely pbranches
A lonely pbranch will keep the enclosing scope alive (up to the function
scope) until the pbranch exits.
@@ -706,74 +778,11 @@ So, in the example above, even though the test will return without waiting for
the pbranch to complete, its stack variables (`var` in this case) will
consume memory until the pbranch returns.
-[[sec-productivity]]
-== Productivity
-
-The core design element of aesop is to make programming of a concurrent
-server easier, so the trade off for memory and runtime performance should
-be worth it. To evaluate this we examine the code complexity of each of
-server implementations.
-
-.Implementation complexity analysis.
-[[table-complex]]
-[cols="3,1,1,1", options="header"]
-|============================
-| Server Implementation | CC | Mod. CC | SLOC
-| aesop | 16 | 11 | 179
-| thread-per-client | 17 | 12 | 182
-| thread-per-client-nb | 17 | 12 | 184
-| thread-per-op | 22 | 17 | 249
-| thread-pool | 32 | 26 | 313
-| event | 28 | 23 | 341
-|============================
-
-<<table-complex>> compares the code complexity of each server
-implementation using McCabe Cyclomatic Complexity (CC) <<McCabe>>,
-Modified McCabe Cyclomatic Complexity (Mod. CC), and Source Lines of Code
-(SLOC). The CC and Mod. CC metrics were measured using the pmccabe tool,
-version 2.6, created by Paul Bame <<Bame>>, while the SLOC metrics were
-measured using the sloccount tool, version 2.26,
-created by David A. Wheeler <<Wheeler>>.
-
-To simplify the comparison, error handling was excluded for all servers
-except for assertions on expected return codes. The protocol definition
-(ie, request and acknowledgement structs) as well as helper functions to
-loop over send and receive were not counted in any of the implementations,
-as these were similar in all four. Also note that the Aesop version
-does not include the Aesop standard library, which
-provides a binding between Aesop and the standard POSIX socket API as part
-of its default functionality. The focus
-of this comparison is on the core logic defining the server implementation.
-
-The Aesop and thread-per-client servers are very similar in terms of
-complexity. The slight increase in complexity for the thread-per-client
-server results from
-the additional function calls needed to create and join threads. The
-nonblocking version of the thread-per-client server (thread-per-client-nb)
-uses two additional lines of code to place each socket into non-blocking
-mode. The remaining code logic needed to manage the nonblocking socket
-calls is implemented in helper functions which are not included in the
-analysis.
-
-The thread-pool and event models are both much more complex than the
-thread-per-client or aesop model. In the case of the thread-pool server,
-this additional complexity arises from not only the queueing and thread
-management logic, but also the event loop which is necessary to detect
-incoming requests and dispatch them to the queue. The event server
-complexity arises from the necessity of dividing servicing routines
-into multiple sub-functions and manually tracking state between
-those functions. An additional complexity of the event
-model which is not captured by these metrics is the fact that control flow
-is not preserved across the processing of a given request. For example,
-servicing a write operation requires 5 disconnected event handlers.
-Although the event model appears less complex than the thread-pool model according to CC and
-Mod. CC, qualitatively it is significantly more challenging to develop.
-
[[sec-compile]]
== Compile Time Performance
Currently aesop imposes some overhead when compiling aesop source. This is
-demomnstrated in a simple micro-benchmark. The test takes an existing aesop
+demonstrated in a simple micro-benchmark. The test takes an existing aesop
source fill with 4200 lines or source and is about 116 KB in size and compiles
it to an object file. The same source file is then renamed to a .c file and
four +#define+ are added which redefine the aesop keywords to nothing.
hooks/post-receive
--
aesop Repository
1
0