Triton-commits
Threads by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
February 2015
- 2 participants
- 63 discussions
26 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via beac38cef731f7875b42ecb63a90150d9bc98a02 (commit)
from 14c8824a5ee66d062e38b11b1f3be3fe5b7e2da3 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit beac38cef731f7875b42ecb63a90150d9bc98a02
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Thu Feb 26 21:11:28 2015 -0500
edits
-----------------------------------------------------------------------
Summary of changes:
.../localstore-report.bib | 9 +++
.../localstore-report.tex | 54 ++++++++++---------
2 files changed, 37 insertions(+), 26 deletions(-)
Diff of changes:
diff --git a/reports/triton-localstore-convergence/localstore-report.bib b/reports/triton-localstore-convergence/localstore-report.bib
index ac6daa1..3e3aded 100644
--- a/reports/triton-localstore-convergence/localstore-report.bib
+++ b/reports/triton-localstore-convergence/localstore-report.bib
@@ -79,3 +79,12 @@ pages={1942--1953},
year={2013},
publisher={VLDB Endowment}
}
+
+@INPROCEEDINGS{6702617,
+author={Soumagne, J. and Kimpe, D. and Zounmevo, J. and Chaarawi, M. and Koziol, Q. and Afsahi, A. and Ross, R.},
+booktitle={2013 IEEE International Conference on Cluster Computing},
+title={Mercury: Enabling remote procedure call for high-performance computing},
+year={2013},
+month={Sept},
+pages={1-8}
+}
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index 84189ce..d4c80b7 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -13,21 +13,22 @@ Report 2.1.6a}
\date{March 30, 2015}
\maketitle
\section{Introduction}
-This report satisfies a portion of the deliverable 2.1.6 to ``Continue
+This report satisfies a portion of the 2.1.6 deliverable to ``Continue
convergence with the Sandia storage system prototype by establishing a
shared API for local storage abstraction, establishing a shared mechanism
for managing client/server communication, and convergence on an overall
fault detection strategy.'' Specifically, it addresses the first portion
of this deliverable: convergence on a shared API for local storage abstraction.
-Both the Triton (ANL) and Sirocco (Sandia) teams have
-developed local storage abstraction prototypes to serve as proofs of concept and
-to explore local storage semantics in previous works. Both teams identified the need for
-for variable-sized, record-oriented access models, atomic primitives, and
-record-granular versioning. We therefore leveraged lessons learned from the initial
-prototypes to produce a new, common local storage component known as the
-Hierarchical Object Storage System (HOSS). HOSS will serve as the storage
-foundation for both the Triton and Sirocco storage systems moving forward.
+The Triton (ANL) and Sirocco (Sandia) teams each developed local storage
+abstraction prototypes to serve as proofs of concept and to explore local
+storage semantics in previous work. We identified common requirements
+for variable-sized, record-oriented access models, atomic primitives,
+and record-granular versioning. We then leveraged lessons learned from
+the initial prototypes to produce a new, common local storage component
+known as the Hierarchical Object Storage System (HOSS). HOSS will serve
+as the storage foundation for both the Triton and Sirocco storage systems
+moving forward.
This report summarizes the architecture of HOSS, describes how it is integrated
into the Triton storage system design, and presents preliminary performance
@@ -91,28 +92,28 @@ tunable log granularity. Each log is stored in a separate userspace file.
\item Direct I/O support: bypass the operating system buffer cache for high
performance storage devices
\item Automatic file and memory alignment: automatically convert unaligned
-access to aligned access for devices that benefit form page alignment
+access to aligned access for devices that benefit from page alignment
\item Use the fallocate(FALLOC\_FL\_PUNCH\_HOLE) system call (where available)
for log garbage collection.
fallocate(FALLOC\_FL\_PUNCH\_HOLE) is an operating system specific feature
-that can be used to reclaim unused disk blocks with fine
-granularity~\cite{fallocate-punch}.
+that can be used to reclaim unused disk blocks without deleting an entire
+file~\cite{fallocate-punch}.
\end{itemize}
The second available Recordstore module is the ``memory'' module, which uses dynamically
allocated memory for storage rather than persistent files. This method is
valuable for debugging, caching, and burst buffer use cases. The memory
-allocation is organized to facilitate potential exploration of anti-caching
+allocation is organized to facilitate the potential exploration of anti-caching
I/O strategies~\cite{debrabant2013anti} in future work.
The third available Recordstore module is the ``Kinetic'' module, which provides
support for the Seagate Kinetic~\cite{kinetic} line of storage
-devices. Kinetic devices bypass the traditional block device abstraction to
-present a native key/value oriented interface.
+devices. Kinetic devices replace the traditional block device abstraction
+with a native key/value interface.
The current family of Recordstore modules demonstrates HOSS's ability to rapidly
support a diverse storage technologies. The architecture
-of HOSS (which separates raw data storage from indexing and metadata) allows all
+of HOSS (which separates raw data storage from indexing and metadata) also allows all
three of these modules to use the same IDB implementation for
atomicity and versioning, thereby simplifying the task of supporting new raw
storage devices.
@@ -132,7 +133,7 @@ database library. Berkeley DB provides transactional updates and range query
capability. The IDB data structures are stored in Berkeley DB using custom
BTree comparison functions for efficiency. Adjacent objects and records
can therefore be retrieved by performing a single BTree search and then
-walking the tree sequentially. This not only reduces search
+walking the tree sequentially. This approach not only reduces search
overhead but also organizes keys in a cache-friendly manner so that database
queries rarely require disk access.
@@ -150,8 +151,8 @@ path in the Triton storage system. HOSS is responsible for local storage on
each server daemon. An additional layer atop HOSS, known as the Replicated
Object Storage Devices (ROSD) presents a similar interface to HOSS, except
that it manages inter-node replication. The ROSD layer must therefore
-interact with both local storage (HOSS) and remote RPC (Mercury)
-sub-components.
+interact with both local storage (HOSS) and remote RPC
+(Mercury)~\cite{6702617} sub-components.
When an application writes data into the storage system, the ROSD layer
forwards it to the appropriate server. The master server applies the write
@@ -175,8 +176,8 @@ are not shown in this figure.
The ROSD layer is also responsible for \emph{pipelining} large I/O transfers.
Read or write operations that are larger than a configurable threshold are
-automatically broken into smaller segments in order to overlap network and
-disk activity while minimizing the required server memory footprint. In this
+automatically broken into smaller segments to overlap network and
+disk activity while minimizing the server's memory footprint. In this
scenario, a single application I/O request may be split into multiple HOSS
local storage operations. These HOSS operations are issued concurrently to the
degree possible using Aesop \emph{pbranch} constructs~\cite{kimpe2012aesop}.
@@ -190,13 +191,14 @@ storage.
\subsection{Current prototype status}
We have developed a prototype version of the Triton storage system that uses
-the HOSS local storage system as described in this document, including both
-replication and pipelining. Based on our experience thus far we are working
+the HOSS local storage system as described in this document, including reads
+and writes with
+replication and pipelining. Based on our experience thus far, we are working
with Sandia to continue to improve the HOSS implementation and expand it to
include more functionality for querying and managing data once it has been
stored in Triton.
-\subsection{Evaluation}
+\subsection{Preliminary Evaluation}
\begin{figure}[t]
@@ -238,8 +240,8 @@ dd if=/dev/zero of=/tmp/1.dat bs=$SIZE count=50000 oflag=direct,dsync
Each configuration was executed and measured for 10 seconds. In the
case of \texttt{dd}, the transfer was halted after 10 seconds by sending a signal to the
-benchmark process, while the Recordstore used
-a custom benchmark utility that automatically halted execution after 10 seconds.
+benchmark process, while the Recordstore
+benchmark automatically halted execution after 10 seconds.
All writes were issued sequentially starting from file offset 0 or record
index 0 (for \texttt{dd} file access and Recordstore object access,
respectively).
hooks/post-receive
--
1
0
26 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 14c8824a5ee66d062e38b11b1f3be3fe5b7e2da3 (commit)
from 330b466a84e34e40e75f13c3940bb6da0a1a2d18 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 14c8824a5ee66d062e38b11b1f3be3fe5b7e2da3
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Thu Feb 26 20:46:29 2015 -0500
editing
-----------------------------------------------------------------------
Summary of changes:
.../localstore-report.tex | 182 ++++++++++----------
1 files changed, 95 insertions(+), 87 deletions(-)
Diff of changes:
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index e07e015..84189ce 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -7,29 +7,31 @@
\usepackage[tight,footnotesize]{subfigure}
\usepackage[pdftex]{graphicx}
\begin{document}
-\title{Convergence on a shared API for local storage abstraction\\Deliverable Report 2.3.6a}
+\title{Convergence on a shared API for local storage abstraction\\Deliverable
+Report 2.1.6a}
\author{Argonne National Laboratory}
\date{March 30, 2015}
\maketitle
\section{Introduction}
-This report statisfies a portion of the deliverable 2.3.6 ``Continue
+This report satisfies a portion of the deliverable 2.1.6 to ``Continue
convergence with the Sandia storage system prototype by establishing a
-shared API for local storage abstraction on establishment a shared mechanism
-for managing client/server communication and convergence on an overall
-fault detection strategy.'' This report specifically addresses the first portion
+shared API for local storage abstraction, establishing a shared mechanism
+for managing client/server communication, and convergence on an overall
+fault detection strategy.'' Specifically, it addresses the first portion
of this deliverable: convergence on a shared API for local storage abstraction.
-Both the Triton (ANL) and Sirocco (Sandia) teams have previously
-developed local storage prototypes to serve as proofs of concept and
-to explore local storage semantics. Both teams identified the need for
+Both the Triton (ANL) and Sirocco (Sandia) teams have
+developed local storage abstraction prototypes to serve as proofs of concept and
+to explore local storage semantics in previous works. Both teams identified the need for
for variable-sized, record-oriented access models, atomic primitives, and
-record-granular versioning. We leveraged lessons learned from the initial
+record-granular versioning. We therefore leveraged lessons learned from the initial
prototypes to produce a new, common local storage component known as the
Hierarchical Object Storage System (HOSS). HOSS will serve as the storage
foundation for both the Triton and Sirocco storage systems moving forward.
-This report summarizes the architecture of HOSS as well as how it is integrated
-into the overall Triton storage system design.
+This report summarizes the architecture of HOSS, describes how it is integrated
+into the Triton storage system design, and presents preliminary performance
+results.
\newpage
\section{HOSS}
@@ -48,7 +50,7 @@ as an \emph{update ID}). HOSS is further divided into two sub-components
known as the Recordstore and IDB. The Recordstore provides bulk storage of
data, while the IDB indexes that data, tracks version numbers, and organizes
data into discrete objects. These components are separated in order to allow
-for rapid introduction of new storage device technologies without replacing
+for rapid adoption of new storage device technologies without replacing
the entire implementation.
When new records are stored in HOSS, the data is first written to the
@@ -62,9 +64,14 @@ persistent storage devices.
\subsection{Recordstore: storage device support}
-The Recordstore component, developed primarily at ANL, is responsible for raw
-data storage, and includes multiple modules to support different back-end storage
-devices. The default module is the ``file'' module, which stores raw data in
+The Recordstore component, developed primarily at ANL, is responsible
+for raw data storage, and includes multiple modules to support
+different back-end storage devices. Data stored in the Recordstore is
+referenced by a unique identifier that is generated by the Recordstore
+itself at write time (in a manner similar to content-addressable
+storage~\cite{Nath08evaluatingthe}).
+
+The default module is the ``file'' module, which stores raw data in
log-structured POSIX files and is therefore compatible with any local
Linux file system. Based on previous experience in local storage performance
engineering~\cite{pvfs-bgp-sc09,carns2010object},
@@ -80,7 +87,7 @@ storage
deliverables) to orchestrate concurrency
\item Log-structured storage: optimize for write-intensive
workloads~\cite{rosenblum1992design} with
-tunable log granularity. Each log is stored in a userspace file.
+tunable log granularity. Each log is stored in a separate userspace file.
\item Direct I/O support: bypass the operating system buffer cache for high
performance storage devices
\item Automatic file and memory alignment: automatically convert unaligned
@@ -92,42 +99,40 @@ that can be used to reclaim unused disk blocks with fine
granularity~\cite{fallocate-punch}.
\end{itemize}
-The second available recordstore module is the ``memory'' module, which uses dynamically
+The second available Recordstore module is the ``memory'' module, which uses dynamically
allocated memory for storage rather than persistent files. This method is
-valuable for debugging as well as caching or burst buffer use cases. The memory
+valuable for debugging, caching, and burst buffer use cases. The memory
allocation is organized to facilitate potential exploration of anti-caching
I/O strategies~\cite{debrabant2013anti} in future work.
-The third available recordstore module is the ``Kinetic'' module, which provides
+The third available Recordstore module is the ``Kinetic'' module, which provides
support for the Seagate Kinetic~\cite{kinetic} line of storage
devices. Kinetic devices bypass the traditional block device abstraction to
-present a native key/value oriented interface to applications.
+present a native key/value oriented interface.
-The current family of recordstore modules demonstrates HOSS's ability to rapidly
-support a diverse storage technologies. We also emphasize that the architecture
+The current family of Recordstore modules demonstrates HOSS's ability to rapidly
+support a diverse storage technologies. The architecture
of HOSS (which separates raw data storage from indexing and metadata) allows all
three of these modules to use the same IDB implementation for
atomicity and versioning, thereby simplifying the task of supporting new raw
-storage devices. Data stored in the recordstore is referenced by a unique
-identifier provided by Recordstore (in a manner similar to
-content-addressible storage~\cite{Nath08evaluatingthe}). This identifier is indexed by the IDB
-component to provide a mapping to local object access.
+storage devices.
+
\subsection{IDB: indexing and versioning}
-The primary roll of the IDB component, primarily developed by Sandia,
-is to track data that has been stored in the Recordstore and organize it into
-a coherent record-based object data model. It records the update ID for each
-record using a sparse data structure so that even single-byte records are stored
-efficiently, and maps I/O operations from logical objects and records to
-abstract chunks of data stored in the Recordstore.
+The IDB component, primarily developed by Sandia,
+tracks data that has been stored in the Recordstore and organize it into
+a coherent record-based object data model. It maps logical record IDs,
+object IDS, and record version numbers to Recordstore IDs using a
+sparse data structure so that even single-byte records are stored
+efficiently.
The IDB is implemented atop the Berkeley DB
database library. Berkeley DB provides transactional updates and range query
capability. The IDB data structures are stored in Berkeley DB using custom
-BTree comparison functions for efficiency. Sequences of objects and records
+BTree comparison functions for efficiency. Adjacent objects and records
can therefore be retrieved by performing a single BTree search and then
-walking the Btree in order from that key. This not only reduces search
+walking the tree sequentially. This not only reduces search
overhead but also organizes keys in a cache-friendly manner so that database
queries rarely require disk access.
@@ -140,28 +145,28 @@ queries rarely require disk access.
\label{fig:replicate}
\end{figure}
-Figure~\ref{fig:replicate} illustrates how HOSS fits into the overall I/O
-path in the Triton storage system. HOSS is reponsible for local storage on
-each server daemon. An additoinal layer atop HOSS, known as the Replicated
+Figure~\ref{fig:replicate} illustrates how HOSS has been integrated into the overall I/O
+path in the Triton storage system. HOSS is responsible for local storage on
+each server daemon. An additional layer atop HOSS, known as the Replicated
Object Storage Devices (ROSD) presents a similar interface to HOSS, except
that it manages inter-node replication. The ROSD layer must therefore
interact with both local storage (HOSS) and remote RPC (Mercury)
-subcomponents.
+sub-components.
When an application writes data into the storage system, the ROSD layer
forwards it to the appropriate server. The master server applies the write
-to HOSS while simulaneously forwarding the replication information to
+to HOSS while simultaneously forwarding the replication information to
secondary servers. Consistency is maintained across servers using the
-record-based versioning capability of HOSS.
+record-based visioning capability of HOSS.
All memory buffers used for data transfer in Figure~\ref{fig:replicate} are
allocated and managed by an explicit buffer management component. The buffer
-manageer constrains the memory resources dedicated to I/O at any
+manager constrains the memory resources dedicated to I/O at any
given time. Each buffer is automatically page-aligned, and the same buffer
-is passed through the ROSD, Mercury, and HOSS layers in order to avoid memory
+is passed through the ROSD, Mercury, and HOSS layers to avoid memory
copies.
-Note that auxilliary internal services within the Triton storage system (such
+Note that auxiliary internal services within the Triton storage system (such
as buffer management, fault detection, and object placement algorithms)
are not shown in this figure.
@@ -187,7 +192,7 @@ storage.
We have developed a prototype version of the Triton storage system that uses
the HOSS local storage system as described in this document, including both
replication and pipelining. Based on our experience thus far we are working
-with Sandia to continue to improve the HOSS implementation and exand it to
+with Sandia to continue to improve the HOSS implementation and expand it to
include more functionality for querying and managing data once it has been
stored in Triton.
@@ -212,13 +217,14 @@ synchronizing each operation to disk.}
\end{figure}
We performed a preliminary evaluation of the Recordstore and HOSS
-components using a streaming write microbenchmark to represent expected
-server-attached storage workloads in an HPC checkpoint scenario. All
+components using a streaming write microbenchmark. The microbenchmark
+reflects expected
+server-attached storage workloads in an HPC checkpoint use case. All
experiments were performed on a Linux 3.16 workstation using the EXT4 file
system and default tuning options. A 240 GiB Intel 730 SSD was used as the
underlying storage device. The write access size was varied from 4 KiB to 4
MiB. All tests were performed using direct I/O (to improve performance for
-high performance storage devices) and each I/O operation was synchronized to
+high performance storage devices), and each I/O operation was synchronized to
disk before being reported as complete (to evaluate durable write
performance).
@@ -230,50 +236,52 @@ comparison. \texttt{dd} was configured as follows:
dd if=/dev/zero of=/tmp/1.dat bs=$SIZE count=50000 oflag=direct,dsync
\end{verbatim}
-Each configuration was executed and measurued for 10 seconds. In the
-case of \texttt{dd}, the transfer was stopped by sending a signal to the
-benchmark process after 10 seconds of elapsed time, while the Recordstore used
-a custom benchmark utility that automatically stopped execution after 10 seconds.
-
+Each configuration was executed and measured for 10 seconds. In the
+case of \texttt{dd}, the transfer was halted after 10 seconds by sending a signal to the
+benchmark process, while the Recordstore used
+a custom benchmark utility that automatically halted execution after 10 seconds.
All writes were issued sequentially starting from file offset 0 or record
index 0 (for \texttt{dd} file access and Recordstore object access,
-respectively). Two recordstore examples are shown: one which issued one
-operation at a time, and one that issued up to 16 concurrent operations using
-Aesop pbranches~\cite{kimpe2012aesop}. The Recordstore configuration without
-concurrency does not perform as well as the raw \texttt{dd} benchmark due to
-threading overhead. The concurrent Recordstore configuration consistently
-exceeds baseline \texttt{dd} performance, however, due to its greater
-ability to saturate the storage device.
+respectively).
+
+Two Recordstore examples are shown in Figure~\ref{fig:rs-perf}: one that
+issued one operation at a time, and one that issued up to 16 concurrent
+operations using Aesop pbranches. Recordstore is optimized for the
+latter scenario, and the underlying Aesop concurrency model allows it to
+saturate available storage bandwidth. As a result, it exceeds the
+baseline \texttt{dd} performance by a wide margin at each access size.
+The non-concurrent workload, in contrast, does not match baseline
+\texttt{dd} performance. We believe that performance for this workload could be
+improved using \texttt{pthread} optimizations to reduce thread latency,
+however.
Figure~\ref{fig:hoss-perf} repeats the same experiment using the HOSS API
-layer atop the Recordstore rather than using the recordstore directly. This
-API layer presents a full object interface with versioning
-and indexing. It's performance trends largely match that of the direct
-Recordstore write benchmark. The most noticable discrepancy is in small (32
-KiB or smaller) write access performance without concurrency. In this
-scenario, the additional Berkeley DB access latency incurs a small
-performance overhead.
-
-We believe that both the Recordstore and HOSS performance could be improved
-via thread tuning and Berkeley DB access optimization, but from these figures
-we observe that both components already perform well for highly
-concurrent workloads. We also note that HOSS is largely agnostic whether the
-workload is sequential or not, due to the fact that the underlying writes in
-the Recordstore component are log-structured. We performed sequential tests
-in these preliminary measurements for simplicity and for straightforward
-comparison to the \texttt{dd} utility. We will continue to optimize both components and
-evaluate their performance for a broader range of workloads.
-
-\section{Conclusion}
-
-The current Triton prototype operating atop the HOSS local storage component
-has demonstrated that we are able to share the same local storage
-implementation as is used by the Sirocco team at Sandia. This will allow us
-to pool research, development, testing, and maintenance effort for this
-component of our respective storage systems moving forward. Preliminary
-performance results also indicate that the HOSS storage component performs
-well for highly concurrent write workloads, and we will continue to evaluate
-its performance in future work.
+layer atop the Recordstore rather than using the Recordstore directly.
+This API layer presents a full object interface with versioning and
+indexing. It's performance trends largely match that of the direct
+Recordstore write benchmark. The most noticeable discrepancy relative to
+the Recordstore performance is in small (32 KiB or smaller) write access
+without concurrency. Berkeley DB access latency is a significant factor
+in response time for small writes, but this effect is mitigated by improved
+utilization under concurrent workloads.
+
+Recordstore and HOSS performance could be improved via thread
+tuning and Berkeley DB access optimization. Both components already
+perform well for highly concurrent workloads, however. We also note that we
+expect to achieve the same bandwidth for non-sequential access as well due to
+the log-structured nature of the Recordstore file module. We focused on
+sequential access patterns in these preliminary experiments for simplicity
+and to facilitate comparison to the sequential \texttt{dd} utility.
+
+\section{Conclusions}
+
+The current Triton prototype operating atop the HOSS local storage
+component has demonstrated that we are able to share the same local
+storage implementation as is used by the Sirocco team at Sandia. This
+will allow us to pool research, development, testing, and maintenance
+efforts for this activity. Preliminary performance results also indicate
+that the HOSS storage component performs well for highly concurrent write
+workloads. We will continue to evaluate its performance in future work.
\bibliographystyle{plain}
\bibliography{localstore-report}
hooks/post-receive
--
1
0
26 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 330b466a84e34e40e75f13c3940bb6da0a1a2d18 (commit)
via 6ae67147597f682781960346b393bf38a94006ce (commit)
via 59e8f04c731a5cc34e42375db68bfb36c66877fb (commit)
from cf43be4262f3e2a38f3e8ff4883f17b0563f1eef (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 330b466a84e34e40e75f13c3940bb6da0a1a2d18
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Thu Feb 26 16:48:23 2015 -0500
bibtex dep in makefile
commit 6ae67147597f682781960346b393bf38a94006ce
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Thu Feb 26 16:48:09 2015 -0500
bib file
commit 59e8f04c731a5cc34e42375db68bfb36c66877fb
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Thu Feb 26 16:47:59 2015 -0500
citations
-----------------------------------------------------------------------
Summary of changes:
reports/triton-localstore-convergence/Makefile | 4 +-
.../localstore-report.bib | 81 ++++++++++++++++++++
.../localstore-report.tex | 39 ++++++----
3 files changed, 107 insertions(+), 17 deletions(-)
create mode 100644 reports/triton-localstore-convergence/localstore-report.bib
Diff of changes:
diff --git a/reports/triton-localstore-convergence/Makefile b/reports/triton-localstore-convergence/Makefile
index d24d573..f4d60a7 100644
--- a/reports/triton-localstore-convergence/Makefile
+++ b/reports/triton-localstore-convergence/Makefile
@@ -1,8 +1,10 @@
all:: localstore-report.pdf
-%.pdf: %.tex
+localstore-report.pdf: localstore-report.tex localstore-report.bib
pdflatex $<
pdflatex $<
+ bibtex localstore-report
+ pdflatex $<
%.draft.pdf: %.pdf
pdftk $< background ../draft.pdf output $@
diff --git a/reports/triton-localstore-convergence/localstore-report.bib b/reports/triton-localstore-convergence/localstore-report.bib
new file mode 100644
index 0000000..ac6daa1
--- /dev/null
+++ b/reports/triton-localstore-convergence/localstore-report.bib
@@ -0,0 +1,81 @@
+@article{shavit1997software,
+ title={Software transactional memory},
+ author={Shavit, Nir and Touitou, Dan},
+ journal={Distributed Computing},
+ volume={10},
+ number={2},
+ pages={99--116},
+ year={1997},
+ publisher={Springer}
+}
+
+@inproceedings{pvfs-bgp-sc09,
+author = {Lang, Samuel and Carns, Philip and Latham, Robert and Ross, Robert and Harms, Kevin and Allcock, William},
+title = {{I/O} performance challenges at leadership scale},
+booktitle = {{SC} '09: Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis},
+year = {2009},
+isbn = {978-1-60558-744-8},
+pages = {1--12},
+location = {Portland, Oregon},
+doi = {http://doi.acm.org/10.1145/1654059.1654100},
+publisher = {ACM},
+address = {New York, NY, USA},
+}
+
+@inproceedings{carns2010object,
+title={Object storage semantics for replicated concurrent-writer file systems},
+author={Carns, P. and Ross, R. and Lang, S.},
+booktitle={Proceedings of 2010 Workshop on Interfaces and Architectures for Scientific Data Storage (IASDS 2010)},
+year={2010},
+organization={IEEE}
+}
+
+@inproceedings{kimpe2012aesop,
+title={AESOP: Expressing Concurrency in High-Performance System Software},
+author={Kimpe, D. and Carns, P. and Harms, K. and Wozniak, J.M. and Lang, S. and Ross, R.},
+booktitle={Proceedings of 7th IEEE International Conference on Networking, Architecture, and Storage (NAS 2012)},
+year={2012}
+}
+
+@article{rosenblum1992design,
+title={The design and implementation of a log-structured file system},
+author={Rosenblum, Mendel and Ousterhout, John K},
+journal={ACM Transactions on Computer Systems (TOCS)},
+volume={10},
+number={1},
+pages={26--52},
+year={1992},
+publisher={ACM}
+}
+
+@misc{kinetic,
+ author = {{Seagate Technology LLC}},
+ title = {{Segate Kinetic}},
+ howpublished = "\url{http://www.seagate.com/Kinetic}"
+}
+
+@misc{fallocate-punch,
+ author = {{Linux man-pages project}},
+ title = {fallocate man page},
+ howpublished =
+ "\url{http://man7.org/linux/man-pages/man2/fallocate.2.html}"
+}
+
+@INPROCEEDINGS{Nath08evaluatingthe,
+author = {Partho Nath and Bhuvan Urgaonkar and Anand Sivasubramaniam},
+title = {Evaluating the usefulness of content addressable storage for high-performance data intensive applications},
+booktitle = {In Proceedings of the 17th High Performance Distributed Computing (HPDC ’08},
+year = {2008},
+publisher = {ACM}
+}
+
+@article{debrabant2013anti,
+title={Anti-caching: A new approach to database management system architecture},
+author={DeBrabant, Justin and Pavlo, Andrew and Tu, Stephen and Stonebraker, Michael and Zdonik, Stan},
+journal={Proceedings of the VLDB Endowment},
+volume={6},
+number={14},
+pages={1942--1953},
+year={2013},
+publisher={VLDB Endowment}
+}
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index e56fa9e..e07e015 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -3,6 +3,7 @@
\usepackage{listings}
\usepackage{appendix}
\usepackage{color}
+\usepackage{url}
\usepackage[tight,footnotesize]{subfigure}
\usepackage[pdftex]{graphicx}
\begin{document}
@@ -56,7 +57,7 @@ data visible to future readers. Updates can therefore proceed with a high
degree of concurrency. The most expensive bulk storage operations need not be
serialized, and the scope of the atomic update is minimized. This strategy
is similar to that employed by most software transactional memory
-architectures \textcolor{red}{TODO: cite}, but we have adapted it for use in
+architectures~\cite{shavit1997software}, but we have adapted it for use in
persistent storage devices.
\subsection{Recordstore: storage device support}
@@ -66,8 +67,8 @@ data storage, and includes multiple modules to support different back-end storag
devices. The default module is the ``file'' module, which stores raw data in
log-structured POSIX files and is therefore compatible with any local
Linux file system. Based on previous experience in local storage performance
-engineering \textcolor{red}{TODO: cite perf challenges at leadership scale, vosd
-paper}, we have implemented the following features in this module:
+engineering~\cite{pvfs-bgp-sc09,carns2010object},
+we have implemented the following features in this module:
\begin{itemize}
\item File descriptor caching: avoid repeatedly opening and closing frequently
@@ -75,28 +76,30 @@ accessed underlying data files
\item Fdatasync() coalescing: coalesce concurrent object flush requests in
order to reduce the number of synchronization operations propagated to
storage
-\item Multithreading: use the Aesop programming model (completed in previous
-deliverables) to orchestrate concurrency \textcolor{red}{TODO: cite}
-\item Log-structured storage: optimize for write-intensive workloads, with
-tunable log granularity \textcolor{red}{TODO: cite}
+\item Multithreading: use the Aesop programming model~\cite{kimpe2012aesop} (completed in previous
+deliverables) to orchestrate concurrency
+\item Log-structured storage: optimize for write-intensive
+workloads~\cite{rosenblum1992design} with
+tunable log granularity. Each log is stored in a userspace file.
\item Direct I/O support: bypass the operating system buffer cache for high
performance storage devices
\item Automatic file and memory alignment: automatically convert unaligned
access to aligned access for devices that benefit form page alignment
\item Use the fallocate(FALLOC\_FL\_PUNCH\_HOLE) system call (where available)
-for log garbage collection \textcolor{red}{TODO: cite}.
+for log garbage collection.
fallocate(FALLOC\_FL\_PUNCH\_HOLE) is an operating system specific feature
-that can be used to reclaim unused disk blocks with fine granularity.
+that can be used to reclaim unused disk blocks with fine
+granularity~\cite{fallocate-punch}.
\end{itemize}
The second available recordstore module is the ``memory'' module, which uses dynamically
allocated memory for storage rather than persistent files. This method is
valuable for debugging as well as caching or burst buffer use cases. The memory
allocation is organized to facilitate potential exploration of anti-caching
-I/O strategies \textcolor{red}{TODO: cite} in future work.
+I/O strategies~\cite{debrabant2013anti} in future work.
The third available recordstore module is the ``Kinetic'' module, which provides
-support for the Seagate Kinetic \textcolor{red}{TODO: cite} line of storage
+support for the Seagate Kinetic~\cite{kinetic} line of storage
devices. Kinetic devices bypass the traditional block device abstraction to
present a native key/value oriented interface to applications.
@@ -105,7 +108,10 @@ support a diverse storage technologies. We also emphasize that the architecture
of HOSS (which separates raw data storage from indexing and metadata) allows all
three of these modules to use the same IDB implementation for
atomicity and versioning, thereby simplifying the task of supporting new raw
-storage devices.
+storage devices. Data stored in the recordstore is referenced by a unique
+identifier provided by Recordstore (in a manner similar to
+content-addressible storage~\cite{Nath08evaluatingthe}). This identifier is indexed by the IDB
+component to provide a mapping to local object access.
\subsection{IDB: indexing and versioning}
@@ -116,7 +122,7 @@ record using a sparse data structure so that even single-byte records are stored
efficiently, and maps I/O operations from logical objects and records to
abstract chunks of data stored in the Recordstore.
-The IDB is implemented atop the Berkeley DB \textcolor{red}{TODO: cite}
+The IDB is implemented atop the Berkeley DB
database library. Berkeley DB provides transactional updates and range query
capability. The IDB data structures are stored in Berkeley DB using custom
BTree comparison functions for efficiency. Sequences of objects and records
@@ -168,8 +174,7 @@ automatically broken into smaller segments in order to overlap network and
disk activity while minimizing the required server memory footprint. In this
scenario, a single application I/O request may be split into multiple HOSS
local storage operations. These HOSS operations are issued concurrently to the
-degree possible using Aesop \emph{pbranch} constructs \textcolor{red}{TODO:
-cite}.
+degree possible using Aesop \emph{pbranch} constructs~\cite{kimpe2012aesop}.
HOSS provides a \emph{group} construct that can be used to combine multiple
I/O operations into a single atomic unit. This functionality could be used
@@ -234,7 +239,7 @@ All writes were issued sequentially starting from file offset 0 or record
index 0 (for \texttt{dd} file access and Recordstore object access,
respectively). Two recordstore examples are shown: one which issued one
operation at a time, and one that issued up to 16 concurrent operations using
-Aesop pbranches~\cite{TODO}. The Recordstore configuration without
+Aesop pbranches~\cite{kimpe2012aesop}. The Recordstore configuration without
concurrency does not perform as well as the raw \texttt{dd} benchmark due to
threading overhead. The concurrent Recordstore configuration consistently
exceeds baseline \texttt{dd} performance, however, due to its greater
@@ -270,4 +275,6 @@ performance results also indicate that the HOSS storage component performs
well for highly concurrent write workloads, and we will continue to evaluate
its performance in future work.
+\bibliographystyle{plain}
+\bibliography{localstore-report}
\end{document}
hooks/post-receive
--
1
0
26 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via cf43be4262f3e2a38f3e8ff4883f17b0563f1eef (commit)
from a76003805c31ab6b81b7003269ef8d04025d1afe (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit cf43be4262f3e2a38f3e8ff4883f17b0563f1eef
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Thu Feb 26 16:11:10 2015 -0500
very rough perf results writup
-----------------------------------------------------------------------
Summary of changes:
.../localstore-report.tex | 61 +++++++++++++++++++-
1 files changed, 58 insertions(+), 3 deletions(-)
Diff of changes:
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index 46c936a..e56fa9e 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -188,6 +188,7 @@ stored in Triton.
\subsection{Evaluation}
+
\begin{figure}[t]
\centering
\subfigure[Recordstore (RS)]{
@@ -202,10 +203,61 @@ stored in Triton.
}
\caption{Streaming write performance at the Recordstore and HOSS component level on an Intel 730 series SSD using direct I/O and
synchronizing each operation to disk.}
+ \label{fig:perf}
\end{figure}
-Figure~\ref{fig:rs-perf} shows something and Figure~\ref{fig:hoss-perf} shows something else. \textcolor{red}{TODO: fill this in.
-Preliminary results. See .txt files for command line details.}
+We performed a preliminary evaluation of the Recordstore and HOSS
+components using a streaming write microbenchmark to represent expected
+server-attached storage workloads in an HPC checkpoint scenario. All
+experiments were performed on a Linux 3.16 workstation using the EXT4 file
+system and default tuning options. A 240 GiB Intel 730 SSD was used as the
+underlying storage device. The write access size was varied from 4 KiB to 4
+MiB. All tests were performed using direct I/O (to improve performance for
+high performance storage devices) and each I/O operation was synchronized to
+disk before being reported as complete (to evaluate durable write
+performance).
+
+The results of these experiments are shown in Figure~\ref{fig:perf}.
+Linux \texttt{dd} command line utility performance is also shown for
+comparison. \texttt{dd} was configured as follows:
+
+\begin{verbatim}
+dd if=/dev/zero of=/tmp/1.dat bs=$SIZE count=50000 oflag=direct,dsync
+\end{verbatim}
+
+Each configuration was executed and measurued for 10 seconds. In the
+case of \texttt{dd}, the transfer was stopped by sending a signal to the
+benchmark process after 10 seconds of elapsed time, while the Recordstore used
+a custom benchmark utility that automatically stopped execution after 10 seconds.
+
+All writes were issued sequentially starting from file offset 0 or record
+index 0 (for \texttt{dd} file access and Recordstore object access,
+respectively). Two recordstore examples are shown: one which issued one
+operation at a time, and one that issued up to 16 concurrent operations using
+Aesop pbranches~\cite{TODO}. The Recordstore configuration without
+concurrency does not perform as well as the raw \texttt{dd} benchmark due to
+threading overhead. The concurrent Recordstore configuration consistently
+exceeds baseline \texttt{dd} performance, however, due to its greater
+ability to saturate the storage device.
+
+Figure~\ref{fig:hoss-perf} repeats the same experiment using the HOSS API
+layer atop the Recordstore rather than using the recordstore directly. This
+API layer presents a full object interface with versioning
+and indexing. It's performance trends largely match that of the direct
+Recordstore write benchmark. The most noticable discrepancy is in small (32
+KiB or smaller) write access performance without concurrency. In this
+scenario, the additional Berkeley DB access latency incurs a small
+performance overhead.
+
+We believe that both the Recordstore and HOSS performance could be improved
+via thread tuning and Berkeley DB access optimization, but from these figures
+we observe that both components already perform well for highly
+concurrent workloads. We also note that HOSS is largely agnostic whether the
+workload is sequential or not, due to the fact that the underlying writes in
+the Recordstore component are log-structured. We performed sequential tests
+in these preliminary measurements for simplicity and for straightforward
+comparison to the \texttt{dd} utility. We will continue to optimize both components and
+evaluate their performance for a broader range of workloads.
\section{Conclusion}
@@ -213,6 +265,9 @@ The current Triton prototype operating atop the HOSS local storage component
has demonstrated that we are able to share the same local storage
implementation as is used by the Sirocco team at Sandia. This will allow us
to pool research, development, testing, and maintenance effort for this
-component of our respective storage systems moving forward.
+component of our respective storage systems moving forward. Preliminary
+performance results also indicate that the HOSS storage component performs
+well for highly concurrent write workloads, and we will continue to evaluate
+its performance in future work.
\end{document}
hooks/post-receive
--
1
0
26 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via a76003805c31ab6b81b7003269ef8d04025d1afe (commit)
from 2a74c490ebf926662e046834c1f30d234b9e37e6 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit a76003805c31ab6b81b7003269ef8d04025d1afe
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Thu Feb 26 15:03:38 2015 -0500
clarify buffer mgmt
-----------------------------------------------------------------------
Summary of changes:
.../localstore-report.tex | 12 ++++++++++--
1 files changed, 10 insertions(+), 2 deletions(-)
Diff of changes:
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index ead5f41..46c936a 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -148,9 +148,17 @@ to HOSS while simulaneously forwarding the replication information to
secondary servers. Consistency is maintained across servers using the
record-based versioning capability of HOSS.
+All memory buffers used for data transfer in Figure~\ref{fig:replicate} are
+allocated and managed by an explicit buffer management component. The buffer
+manageer constrains the memory resources dedicated to I/O at any
+given time. Each buffer is automatically page-aligned, and the same buffer
+is passed through the ROSD, Mercury, and HOSS layers in order to avoid memory
+copies.
+
Note that auxilliary internal services within the Triton storage system (such
-as fault detection and object placement algorithms) are not shown in this
-figure.
+as buffer management, fault detection, and object placement algorithms)
+are not shown in this figure.
+
\subsection{Pipelining}
hooks/post-receive
--
1
0
branch, shared-lib, updated. 40fd462d6268eb9e645ec61b051bec81f19d5c05
by noreply@mcs.anl.gov 24 Feb '15
by noreply@mcs.anl.gov 24 Feb '15
24 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, shared-lib has been updated
via 40fd462d6268eb9e645ec61b051bec81f19d5c05 (commit)
via 3b3c8bcdbc726d12d3b14415834e5a4d6ce59709 (commit)
via 3cb24f34e69ef4d96f4439af00b56fe948e52c4e (commit)
via 30e1aa3cf62bc0e086b5d4104a4dcefe578e5faa (commit)
via 17b1d53430bc3432dc1580684046e46d5abd9c06 (commit)
via ab5ccb1ecb433d82597fe597fd3a15e5130413f1 (commit)
via fdb3498a82ba5d36e4ac691fb5a5574d26cb525a (commit)
via b7c7e0b9eaf3a147359f933218ab0a3bc1997e6d (commit)
via 2c4c0cd0c65787353fdb9cc9de9ece9e221c3151 (commit)
via a02a9520730f28b2afa2c1a6740899cfadc8f4ba (commit)
via d65e5e20c0eab9eb43c5d5dfbd99275b0f029c83 (commit)
via 2407d98ffe72d250353721c3a85d672e41c7170d (commit)
via 8ff2109e3ce2812da288af2eefaa5ea18823666f (commit)
via f9a1714b96f3f3ba635d1d37942e15da99d964fc (commit)
via c8004907838a9cdca00615126c2ea63256f42bf0 (commit)
via a3d013ebaac4c89f4d90a3d92c59ece37c9b0a08 (commit)
via 8d1e41a26ebcf843709fe02eb553fd188c00220a (commit)
via 08abed59c94b6b5da805e2b54496ae09942e25cd (commit)
via 4124ed82cb7866bb28b53c33d221bdc53b927ea6 (commit)
via b2cf7bffe6bcaf4387b5548498ac6f193415aad1 (commit)
via 8128c05ae72e45e0f80f888f956356a70fba2e34 (commit)
via 05f4efaa673317dc05262d05c3419c0ecf132976 (commit)
via 1ec468ad7afc2383d6c2f12350fba7965c073abe (commit)
via 37c1946e2803d9e1f0421886b84d400b5a7d28b0 (commit)
via 95f6f394f7d4bde54484a658610457ba46c1f1ef (commit)
via b38c2a1349f63efda356eb4b071bbdedae45478b (commit)
via 98c2ec7e089542bad5b91c262b7836c25839e3bf (commit)
via 1c8f9ce048812172a1d5c8fc9e27e54ccec26afc (commit)
via 5874724ffbdd0f6b74f7f3e1aa6f309f04790c3f (commit)
via 0b78b60a5bd241ade8d21d8298711a24ea5a9d48 (commit)
via 64e70ab86b099ba4185482e12f1d99c188517544 (commit)
via 243858ee26b0f5b216ebca617715fc996b21d446 (commit)
via 0381c36fd22d2d1d8c28e54ab6ae8ef1588c2320 (commit)
via dec2f5f70b2f5ed53d323a629f6cdec77ebce63b (commit)
via 5cb6783f2244e392b589a3a365e6f4329d325bb3 (commit)
via ae6bba9cca9e7bfc9cbe55d0644836681d10e078 (commit)
from a7bfff430e5fd4956068bec02f3304e914e8c67a (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 40fd462d6268eb9e645ec61b051bec81f19d5c05
Author: John Jenkins <jenkins(a)mcs.anl.gov>
Date: Tue Feb 24 16:09:51 2015 -0600
add tests for asg conditional operations
(just ASG_COND_ALL for now)
commit 3b3c8bcdbc726d12d3b14415834e5a4d6ce59709
Author: John Jenkins <jenkins(a)mcs.anl.gov>
Date: Tue Feb 24 15:59:58 2015 -0600
name change in swig
commit 3cb24f34e69ef4d96f4439af00b56fe948e52c4e
Author: John Jenkins <jenkins(a)mcs.anl.gov>
Date: Mon Feb 23 09:34:21 2015 -0600
header generation
commit 30e1aa3cf62bc0e086b5d4104a4dcefe578e5faa
Merge: a7bfff430e5fd4956068bec02f3304e914e8c67a 17b1d53430bc3432dc1580684046e46d5abd9c06
Author: John Jenkins <jenkins(a)mcs.anl.gov>
Date: Mon Feb 23 08:57:05 2015 -0600
Merge remote-tracking branch 'origin/trac-323-sos' into shared-lib
-----------------------------------------------------------------------
Summary of changes:
code/bindings/asg.i | 3 +-
code/include/asg.h | 262 +++----
code/src/admin-tools/triton-cp.ae | 32 +-
code/src/admin-tools/triton-touch.ae | 4 +-
code/src/asg/asg-internal.ae | 122 ++--
code/src/asg/asg-internal.hae | 49 +-
code/src/asg/asg.c | 53 ++-
code/src/common/Makefile.subdir | 3 +-
code/src/common/triton-error.spec | 1 +
code/src/replicated-osd/rosd-create.ae | 17 +-
code/src/replicated-osd/rosd-read.ae | 793 ++++++++++++++------
code/src/replicated-osd/rosd-reset.ae | 8 -
code/src/replicated-osd/rosd-write.ae | 114 ++-
code/src/replicated-osd/rosd.ae | 2 +-
code/src/replicated-osd/rosd.hae | 38 +-
code/src/server/triton-server.ae | 66 ++-
code/src/zeroconf/zeroconf.c | 12 +-
code/tests/Makefile.subdir | 5 +-
code/tests/asg/test-asg-cond.c | 248 ++++++
.../asg/{test-asg-simple.sh => test-asg-cond.sh} | 2 +-
code/tests/asg/test-asg-simple.c | 144 ++++-
code/tests/placement/test-placement.ae | 6 +-
code/tests/test-util.sh | 3 +-
23 files changed, 1444 insertions(+), 543 deletions(-)
create mode 100644 code/tests/asg/test-asg-cond.c
copy code/tests/asg/{test-asg-simple.sh => test-asg-cond.sh} (91%)
Diff of changes:
diff --git a/code/bindings/asg.i b/code/bindings/asg.i
index 5d5adef..a605ec3 100644
--- a/code/bindings/asg.i
+++ b/code/bindings/asg.i
@@ -75,7 +75,8 @@ static asg_instance_t asg_instance_value(asg_instance_t *a) {
/* get rid of ugly asg.asg_* functions in the java code */
%rename(initialize) asg_initialize;
%rename(finalize) asg_finalize;
-%rename(read) asg_read;
+%rename(read_one) asg_read_one;
+%rename(read_sequence) asg_read_sequence;
%rename(write) asg_write;
%rename(punch) asg_punch;
%rename(reset) asg_reset;
diff --git a/code/include/asg.h b/code/include/asg.h
index 55b10f0..f73f303 100644
--- a/code/include/asg.h
+++ b/code/include/asg.h
@@ -4,14 +4,27 @@
* storage system. It presents a hierarchy consisting of containers,
* objects, forks, and records in which data can be stored.
*
- * Revised following meeting on March 17, 2014 to accommodate the following
- * proposed changes:
+ * 1.1.0: Revised following meeting on March 17, 2014 to accommodate the
+ * following proposed changes:
* - switch from "version" to "update id" terminology
* - add probe parameters to asg_read call to support read+probe semantics
* in a single API operation
* .
- * current API level is 1.1.0
*
+ * 1.2.0: Revised in January 2015 to clean up the following issues:
+ * - update Doxygen comments for info_t structs to correctly describe _count
+ * fields
+ * - removed unused asg_offset_t from public API
+ * - changed name of data argument in write() to buf for consistency with
+ * read()
+ * - refactored asg_read() into asg_read_one() and asg_read_sequence()
+ * - clarified and reordered asg_write() arguments
+ * - renamed asg_punch() arguments for clarity
+ * - remove unecessary recordlen argument to asg_reset()
+ * .
+ *
+ * Current API level is 1.2.0
+ *
* @{
*/
@@ -19,8 +32,8 @@
* Declarations for the ASG object storage API.
*/
-#ifndef TRITON_ASG_H
-#define TRITON_ASG_H
+#ifndef __ASG_H
+#define __ASG_H
#include <stdint.h>
#include <string.h>
@@ -39,7 +52,7 @@ extern "C"
/** Minor version number of API
* @note update on minor changes that break API compatibility
*/
-#define ASG_VERSION_MINOR 1
+#define ASG_VERSION_MINOR 2
/** Sub version number of API
* @note update on do not break API compatibility
@@ -95,11 +108,6 @@ typedef uint64_t asg_object_id_t;
typedef uint64_t asg_container_id_t;
/**
- * @todo asg_offset_t seems to be unused, can we delete it?
- */
-typedef uint64_t asg_offset_t;
-
-/**
* Size type, used to indicate the size of records, buffers, etc.
*/
typedef uint64_t asg_size_t;
@@ -153,9 +161,13 @@ typedef uint64_t asg_update_id_t;
typedef enum
{
ASG_SUCCESS = 0, /**< operation succeeded */
- ASG_ERR_UPDATE_ID = 1, /**< operation failed due to conditional test */
- ASG_ERR_LOCATION = 2, /**< operation failed due to invalid location */
- ASG_ERR_OTHER = 666 /**< all other failure modes */
+ ASG_ERR_UPDATE_ID = 1, /**< failed due to conditional test */
+ ASG_ERR_LOCATION = 2, /**< invalid location */
+ ASG_ERR_NOENT = 3, /**< entity does not exist */
+ ASG_ERR_BUFFER_SMALL = 4, /**< buffer too small to hold data */
+ ASG_ERR_RECORDLEN_INVAL = 5, /**< invalid record length */
+
+ ASG_ERR_OTHER = 54321 /**< all other failure modes */
} asg_ret_t;
@@ -193,14 +205,14 @@ typedef enum
* existing update_id numbers in the range. */
ASG_AUTO_UPDATE_ID = 0x0004,
- /** @todo What does this flag do?
- */
- ASG_WRITE_FIXED_DATA = 0x0008,
-
} asg_flags_t;
/** Describes 1 or more records that have contiguous record ID values and
* identical lengths.
+ *
+ * @note asg_record_info_t may describe a sequence of records that all have
+ * the same size but more than one update_id. In that case
+ * record_update_id will be set to ASG_UPDATE_ID_MIXED.
*/
typedef struct
{
@@ -213,17 +225,12 @@ typedef struct
/**
* Describes 1 or more containers that have contiguous container ID values
* and identical object counts.
- *
- * @note It seems unlikely that a sequence of containers will all have the
- * same value for object_count. This capability of describing a sequence of
- * containers with a single container_info_t is mainly here to be consistent
- * with the record_info_t struct.
*/
typedef struct
{
asg_container_id_t container_id; /**< ID of first container in the sequence */
asg_size_t seq_len; /**< Number of containers in sequence */
- asg_size_t object_count; /**< Number of objects within each container */
+ asg_size_t object_count; /**< Total number of objects in the seq_len containers */
} asg_container_info_t;
/**
@@ -234,7 +241,7 @@ typedef struct
{
asg_container_id_t object_id; /**< ID of first object in the sequence */
asg_size_t seq_len; /**< Number of objects in sequence */
- asg_size_t fork_count; /**< Number of forks within each object */
+ asg_size_t fork_count; /**< Total number of forks in the seq_len objects */
} asg_object_info_t;
/**
@@ -245,7 +252,7 @@ typedef struct
{
asg_fork_id_t fork_id; /**< ID of first fork in the sequence */
asg_size_t seq_len; /**< Number of forks in sequence */
- asg_size_t record_count; /**< Number of records within each fork */
+ asg_size_t record_count; /**< Total number of records in the seq_len forks */
} asg_fork_info_t;
/* ========================================================================*/
@@ -274,102 +281,113 @@ int asg_finalize (
);
/**
- * Retrieve data from one or more contiguous records in a fork. The records
- * need not be the same size or update ID number, but they must be contiguous
- * in terms of their record IDs.
+ * Retrieve data from one record.
*
* @param[in] instance Interface instance
* @param[in] location Storage system location
* @param[in] container
* @param[in] object
* @param[in] fork
- * @param[in] start_record First record in the contigous sequence to be read
- * @param[in] recordcount Number of records to read
+ * @param[in] record
* @param[in] flags Flags to modify behavior or specify conditional modes
* @param[in] update_id_condition Update ID value to use for comparison
* tests in conditional mode
- * @param[out] buf Data payload read from disk. May contain data from
- * multiple records serialized into a single byte stream.
- * @param[in] bufsize Size of buf memory provided by caller (bytes)
- * @param[out] transferred Number of bytes read into buf
- * @param[out] record_buf Array describing records read from storage system.
- * @param[in,out] record_buf_transferred Number of record_buf elements
- * provided by the caller on input, and the number of record_buf elements
- * actually filled in by the ASG API on output.
- * @param[out] update_id_info indicates the update ID value of the records
- * that were read from the storage system. Will be set to
- * ASG_UPDATE_ID_MIXED if multiple records were read and they do not all
- * share the same record ID.
+ * @param[out] buf Data payload read from disk
+ * @param[in] buf_size Size of buf memory provided by caller (bytes)
+ * @param[out] record_info Description of record that was read
+ *
+ * @note The actual size of the record is contained in the record_info
+ * output struct.
*
- * @todo Change name of record_buf_transferred parameter if it is really an
- * in,out parameter and not just output?
- * @todo The update_id_info parameter seems redundant: this same information
- * is also present in the record_buf argument.
- * @todo Document the failure modes. What if the records (or some of them)
- * don't exist? What if bufsize or record_buf_transferred is not large
- * enough to hold the payload or record descriptions?
- * @todo Which (if any) out arguments can be set to NULL if you don't care to
- * have them filled in? Possibilities include udpate_id_info,
- * record_buf_transferred, and record_buf.
- */
-int asg_read (
+ * @return
+ * - ASG_SUCCESS on success
+ * - ASG_ERR_NOENT if specified record doesn't exist
+ * - ASG_ERR_BUFFER_SMALL if buf_size isn't large enough to hold record
+ * .
+ */
+int asg_read_one (
asg_instance_t instance,
asg_location_t location,
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
+ asg_record_id_t record,
asg_flags_t flags,
asg_update_id_t update_id_condition,
void * buf,
- size_t bufsize,
- asg_size_t * transferred,
- asg_record_info_t * record_buf,
- asg_size_t * record_buf_transferred,
- asg_update_id_t * update_id_info);
+ size_t buf_size,
+ asg_record_info_t * record_info);
+
+
+
+/**
+ * Retrieve data from a set of records that have sequential record IDs and
+ * identical record lengths.
+ *
+ * @param[in] instance Interface instance
+ * @param[in] location Storage system location
+ * @param[in] container
+ * @param[in] object
+ * @param[in] fork
+ * @param[in] record_start First record in the contigous sequence to be read
+ * @param[in] record_count Number of records to read
+ * @param[in] record_expected_length Expected length of each record in sequence
+ * @param[in] flags Flags to modify behavior or specify conditional modes
+ * @param[in] update_id_condition Update ID value to use for comparison
+ * tests in conditional mode
+ * @param[out] buf Data payload read from disk. Must be be big enough to
+ * hold up to (record_count*record_expected_length) bytes
+ * @param[out] record_info Description of records that were read
+ *
+ * @note asg_read_sequence() may produce a "short" read if record_start
+ * exists but there are not record_count sequential records available. In
+ * this case the actual number of records found is indicated by
+ * record_info->seq_len.
+ *
+ * @return
+ * - ASG_SUCCESS on success
+ * - ASG_ERR_NOENT if specified record_start doesn't exist
+ * - ASG_ERR_RECORDLEN_INVAL if the sequence contains a record with a length
+ * that doesn't match record_expected_length
+ * .
+ */
+int asg_read_sequence (
+ asg_instance_t instance,
+ asg_location_t location,
+ asg_container_id_t container,
+ asg_object_id_t object,
+ asg_fork_id_t fork,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
+ asg_size_t record_expected_length,
+ asg_flags_t flags,
+ asg_update_id_t update_id_condition,
+ void * buf,
+ asg_record_info_t * record_info);
/**
* Write data to one or more contiguous records in a fork. The records
- * @e must all be the same size and update ID number.
+ * must all be the same size and update ID number.
*
* @param[in] instance Interface instance
* @param[in] location Storage system location
* @param[in] container
* @param[in] object
* @param[in] fork
- * @param[in] start_record First record in the contigous sequence to be
+ * @param[in] record_start First record in the contigous sequence to be
* written
- * @param[in] recordcount Number of records to be written
- * @param[in] recordlen Size of each record to be written
+ * @param[in] record_count Number of records to be written
+ * @param[in] record_length Size of each record to be written
* @param[in] flags Flags to modify behavior or specify conditional modes
* @param[in] update_id_condition Update ID value to use for comparison
* tests in conditional mode
+ * @param[out] buf Data payload to be written. May contain data from
+ * multiple records serialized into a single byte stream. The size of the
+ * buffer in bytes must be record_count*record_length.
+ * @param[out] records_written The number of records
+ * written (in case of short write, as in ASG_COND_UNTIL)
* @param[out] new_update_id Update ID value stored on disk; this is only
* filled in when using ASG_AUTO_UPDATE_ID flag.
- * @param[out] data Data payload to be written. May contain data from
- * multiple records serialized into a single byte stream. The size of the
- * data buffer in bytes must be recordcount*recordlen.
- * @param[out] transferred The number of records written (in case of short
- * write, as in ASG_COND_UNTIL)
- * @param[in] bufsize Size of buf memory provided by caller (bytes)
- * @param[out] transferred Number of bytes read into buf
- * @param[out] record_buf Array describing records read from storage system.
- * @param[in,out] record_buf_transferred Number of record_buf elements
- * provided by the caller on input, and the number of record_buf elements
- * actually filled in by the ASG API on output.
- * @param[out] update_id_info indicates the update ID value of the records
- * that were read from the storage system. Will be set to
- * ASG_UPDATE_ID_MIXED if multiple records were read and they do not all
- * share the same record ID.
- *
- * @todo Why is the data payload argument called "buf" in the read function
- * but "data" in the write function?
- * @todo Confirm the units of the "transferred" argument and then adjust
- * name accordingly for clarity. Here I assume that it is the number of
- * records successfully written, in which case "records_transferred" would
- * be a more helpful name.
- *
*/
int asg_write (
asg_instance_t instance,
@@ -377,14 +395,14 @@ int asg_write (
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
- asg_size_t recordlen,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
+ asg_size_t record_length,
asg_flags_t flags,
asg_update_id_t update_id_condition,
asg_update_id_t * new_update_id,
- const void * data,
- asg_size_t * transferred);
+ const void * buf,
+ asg_size_t * records_written);
/**
* Punch is exactly like write but writes zero length records.
@@ -394,21 +412,15 @@ int asg_write (
* @param[in] container
* @param[in] object
* @param[in] fork
- * @param[in] start_record First record in the contigous sequence to be
+ * @param[in] record_start First record in the contigous sequence to be
* punched
- * @param[in] recordcount Number of records to be punched
+ * @param[in] record_count Number of records to be punched
* @param[in] flags Flags to modify behavior or specify conditional modes
* @param[in] update_id_condition Update ID value to use for comparison
* tests in conditional mode
* @param[out] new_update_id Update ID value stored on disk; this is only
* filled in when using ASG_AUTO_UPDATE_ID flag.
- * @param[out] transferred The number of records punched
- *
- * @todo Confirm the units of the "transferred" argument and then adjust
- * name accordingly for clarity. Here I assume that it is the number of
- * records successfully punched, in which case "records_punched" would
- * be a more helpful name.
- *
+ * @param[out] records_punched The number of records punched
*/
int asg_punch (
asg_instance_t instance,
@@ -416,12 +428,12 @@ int asg_punch (
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
asg_flags_t flags,
asg_update_id_t update_id_condition,
asg_update_id_t * new_update_id,
- asg_size_t * transferred);
+ asg_size_t * records_punched);
/**
@@ -436,14 +448,13 @@ int asg_punch (
* @param[in] container
* @param[in] object
* @param[in] fork
- * @param[in] start_record First record in the contigous sequence to be
+ * @param[in] record_start First record in the contigous sequence to be
* reset
- * @param[in] recordcount Number of records to be reset
- * @param[in] recordlen
+ * @param[in] record_count Number of records to be reset
* @param[in] flags Flags to modify behavior or specify conditional modes
* @param[in] update_id_condition Update ID value to use for comparison
* tests in conditional mode
- * @param[out] reset_count The number of records reset
+ * @param[out] records_reset The number of records reset
*
* @note Reset can be used to reset more than just a specific set of
* records:
@@ -452,8 +463,6 @@ int asg_punch (
* - Passing ASG_OBJECT_NULL resets all objects in the container.
* - Passing ASG_CONTAINER_NULL resets all containers in the storage system.
*
- * @todo what does recordlen do in this context? It seems like an entity
- * that has been reset to its initial state will always have a length of 0.
*/
int asg_reset (
asg_instance_t instance,
@@ -461,12 +470,11 @@ int asg_reset (
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
- asg_size_t recordlen,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
asg_flags_t flags,
asg_update_id_t update_id_condition,
- asg_size_t * reset_count);
+ asg_size_t * records_reset);
/**
@@ -488,11 +496,6 @@ int asg_reset (
* be passed to 'start' in subsequent calls). If no more items are
* available then it will be set to ASG_CONTAINER_NULL.
*
- * @todo Is there any faster mechanism to use if the caller does not wish to
- * retrieve the object count for each container in the system?
- * @todo this function separates the incoming size of the array and outgoing
- * size into separate arguments, unlike earlier functions in this header.
- * Which convention do we want to use? We should be consistent.
*/
int asg_probe_system (
asg_instance_t instance,
@@ -523,11 +526,6 @@ int asg_probe_system (
* be passed to 'start' in subsequent calls). If no more items are
* available then it will be set to ASG_OBJECT_NULL.
*
- * @todo Is there any faster mechanism to use if the caller does not wish to
- * retrieve the fork count for each object in a container?
- * @todo this function separates the incoming size of the array and outgoing
- * size into separate arguments, unlike earlier functions in this header.
- * Which convention do we want to use? We should be consistent.
*/
int asg_probe_container (
asg_instance_t instance,
@@ -560,11 +558,6 @@ int asg_probe_container (
* be passed to 'start' in subsequent calls). If no more items are
* available then it will be set to ASG_FORK_NULL.
*
- * @todo Is there any faster mechanism to use if the caller does not wish to
- * retrieve the record count for each fork in an object?
- * @todo this function separates the incoming size of the array and outgoing
- * size into separate arguments, unlike earlier functions in this header.
- * Which convention do we want to use? We should be consistent.
*/
int asg_probe_object (
asg_instance_t instance,
@@ -599,11 +592,6 @@ int asg_probe_object (
* be passed to 'start' in subsequent calls). If no more items are
* available then it will be set to ASG_RECORD_NULL.
*
- * @todo Is there any faster mechanism to use if the caller does not wish to
- * retrieve the record length for each record in a fork?
- * @todo this function separates the incoming size of the array and outgoing
- * size into separate arguments, unlike earlier functions in this header.
- * Which convention do we want to use? We should be consistent.
*/
int asg_probe_fork (
asg_instance_t instance,
@@ -621,6 +609,6 @@ int asg_probe_fork (
}
#endif
-#endif
+#endif /* __ASG_H */
/* @} */
diff --git a/code/src/admin-tools/triton-cp.ae b/code/src/admin-tools/triton-cp.ae
index 5595360..6e622a3 100644
--- a/code/src/admin-tools/triton-cp.ae
+++ b/code/src/admin-tools/triton-cp.ae
@@ -18,6 +18,7 @@
/* TODO: make this configurable */
#define BUFFER_SZ (32*1024*1024)
+#define DEFAULT_CONTAINER 1
enum obj_ref_type
{
@@ -246,7 +247,7 @@ static __blocking int tc_open(struct obj_ref* ref, int create_if_needed)
if(create_if_needed)
{
/* TODO: make replication factor configurable */
- tret = remote_triton_rpc_rosd_create(0, ref->u.triton.oid,
+ tret = remote_triton_rpc_rosd_create(DEFAULT_CONTAINER, ref->u.triton.oid,
replication_factor, 0);
if(triton_error_equal(tret, TRITON_ERR_EXIST))
{
@@ -276,11 +277,7 @@ static __blocking int tc_read(struct obj_ref* ref, char* buffer, int size)
{
int ret;
triton_ret_t tret;
- uint64_t txn_nr;
- int64_t out_size;
- struct rosd_record_info *rinfo = NULL;
- uint64_t num_record_infos;
-
+ struct rosd_record_info rinfo;
if(ref->type == TYPE_POSIX)
{
@@ -297,31 +294,26 @@ static __blocking int tc_read(struct obj_ref* ref, char* buffer, int size)
}
else if(ref->type == TYPE_TRITON)
{
- rinfo = malloc((size_t)size * sizeof(*rinfo));
- if (!rinfo) return -1;
-
- tret = remote_triton_rpc_rosd_read(
- 0,
+ tret = remote_triton_rpc_rosd_read_sequence(
+ DEFAULT_CONTAINER,
ref->u.triton.oid,
0,
ref->u.triton.offset,
size,
+ 1,
+ 0,
0,
buffer,
- size,
- &txn_nr,
- &out_size,
- rinfo,
- &num_record_infos);
+ &rinfo);
if(triton_is_error(tret))
{
- triton_error_print(tret, "remote_triton_rpc_rosd_read()");
+ triton_error_print(tret, "remote_triton_rpc_rosd_read_sequence()");
triton_error_destroy(tret);
return(-1);
}
triton_error_destroy(tret);
- ref->u.triton.offset += out_size;
- ret = out_size;
+ ref->u.triton.offset += rinfo.num_records;
+ ret = rinfo.num_records;
}
else
{
@@ -352,7 +344,7 @@ static __blocking int tc_write(struct obj_ref* ref, const char* buffer, int size
else if(ref->type == TYPE_TRITON)
{
tret = remote_triton_rpc_rosd_write(
- 0,
+ DEFAULT_CONTAINER,
ref->u.triton.oid,
0,
ref->u.triton.offset,
diff --git a/code/src/admin-tools/triton-touch.ae b/code/src/admin-tools/triton-touch.ae
index 8f82713..3dd2dd9 100644
--- a/code/src/admin-tools/triton-touch.ae
+++ b/code/src/admin-tools/triton-touch.ae
@@ -13,6 +13,8 @@
#include "src/system-state/system-state.hae"
#include "src/replicated-osd/rosd.hae"
+#define DEFAULT_CONTAINER 1
+
__blocking int aesop_main(int argc, char **argv)
{
triton_ret_t tret;
@@ -81,7 +83,7 @@ __blocking int aesop_main(int argc, char **argv)
}
/* create object */
- tret = remote_triton_rpc_rosd_create(0, oid, replication_factor, 0);
+ tret = remote_triton_rpc_rosd_create(DEFAULT_CONTAINER, oid, replication_factor, 0);
if(triton_is_error(tret) && !triton_error_equal(tret, TRITON_ERR_EXIST))
{
triton_error_print(tret, "remote_triton_rpc_rosd_create()");
diff --git a/code/src/asg/asg-internal.ae b/code/src/asg/asg-internal.ae
index af49791..c11cbdf 100644
--- a/code/src/asg/asg-internal.ae
+++ b/code/src/asg/asg-internal.ae
@@ -155,68 +155,104 @@ __blocking int asg_i_write (
return(ASG_SUCCESS);
}
-__blocking int asg_i_read (
+__blocking int asg_i_read_sequence (
asg_instance_t instance,
asg_location_t location,
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
+ asg_size_t record_expected_len,
asg_flags_t flags,
- asg_update_id_t version_condition,
- void * buf,
- size_t bufsize,
- asg_size_t * transferred,
- asg_record_info_t * record_buf,
- asg_size_t * record_buf_transferred,
- asg_update_id_t * version_info)
+ asg_update_id_t update_id_condition,
+ const void * buf,
+ asg_record_info_t * record_info)
{
triton_ret_t tret;
- int64_t size;
int ret = ASG_SUCCESS;
- struct rosd_record_info *rinfo = NULL;
- asg_size_t i;
+ struct rosd_record_info rinfo;
if (instance!=1) return ASG_ERR_OTHER;
- size = MIN(bufsize, recordcount);
-
- rinfo = malloc(recordcount * sizeof(*rinfo));
- if (!rinfo) return ASG_ERR_OTHER;
-
- tret = remote_triton_rpc_rosd_read(
+ tret = remote_triton_rpc_rosd_read_sequence(
(uint64_t)container,
(uint64_t)object,
(uint64_t)fork,
- (uint64_t)start_record,
- (uint64_t)recordcount,
- (uint32_t) flags,
+ (uint64_t)record_start,
+ (uint64_t)record_count,
+ (uint64_t)record_expected_len,
+ (uint32_t)flags,
+ (uint64_t)update_id_condition,
buf,
- (uint64_t)bufsize,
- (uint64_t*)version_info,
- (uint64_t*)transferred,
- rinfo,
- (uint64_t*)record_buf_transferred);
+ &rinfo);
if (triton_is_error(tret)) {
- triton_error_print(tret, "remote_triton_rpc_rosd_read ");
+ /* TODO: do we really want _print fns in the library? */
+ triton_error_print(tret, "remote_triton_rpc_rosd_read_sequence");
ret = ASG_ERR_OTHER;
}
-
- /* manually copy over record metadata (really, this should be a memcpy) */
- for (i = 0; i < *record_buf_transferred; i++) {
- record_buf[i].record_id = rinfo[i].record_id;
- record_buf[i].seq_len = rinfo[i].num_records;
- record_buf[i].record_len = rinfo[i].record_len;
- record_buf[i].record_update_id = rinfo[i].record_update_id;
+ else
+ {
+ /* manually copy over record metadata (really, this should be a memcpy) */
+ record_info->record_id = rinfo.record_id;
+ record_info->seq_len = rinfo.num_records;
+ record_info->record_len = rinfo.record_len;
+ record_info->record_update_id = rinfo.record_update_id;
}
- free(rinfo);
triton_error_destroy(tret);
return ret;
}
+__blocking int asg_i_read_one (
+ asg_instance_t instance,
+ asg_location_t location,
+ asg_container_id_t container,
+ asg_object_id_t object,
+ asg_fork_id_t fork,
+ asg_record_id_t record,
+ asg_flags_t flags,
+ asg_update_id_t update_id_condition,
+ void * buf,
+ size_t buf_size,
+ asg_record_info_t * record_info)
+{
+ triton_ret_t tret;
+ uint64_t record_len;
+ uint64_t record_update_id;
+ int ret;
+
+ tret = remote_triton_rpc_rosd_read_one(
+ container,
+ object,
+ fork,
+ record,
+ flags,
+ update_id_condition,
+ buf,
+ buf_size,
+ &record_len,
+ &record_update_id);
+ if (triton_error_equal(tret, TRITON_ERR_NOENT))
+ ret = ASG_ERR_NOENT;
+ else if (triton_error_equal(tret, TRITON_ERR_RECV_TOO_SMALL))
+ ret = ASG_ERR_BUFFER_SMALL;
+ else if (triton_error_equal(tret, TRITON_ERR_CONDITIONAL))
+ ret = ASG_ERR_UPDATE_ID;
+ else if (triton_is_error(tret))
+ ret = ASG_ERR_OTHER;
+ else {
+ ret = ASG_SUCCESS;
+ record_info->record_id = record;
+ record_info->seq_len = 1;
+ record_info->record_len = record_len;
+ record_info->record_update_id = record_update_id;
+ }
+ triton_error_destroy(tret);
+ return ret;
+}
+
__blocking int asg_i_punch (
asg_instance_t instance,
asg_location_t location,
@@ -273,12 +309,11 @@ __blocking int asg_i_reset (
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
- asg_size_t recordlen,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
asg_flags_t flags,
asg_update_id_t update_id_condition,
- asg_size_t * reset_count)
+ asg_size_t * records_reset)
{
triton_ret_t tret;
uint64_t count;
@@ -301,9 +336,8 @@ __blocking int asg_i_reset (
container,
object,
fork,
- start_record,
- recordcount,
- recordlen,
+ record_start,
+ record_count,
flags,
update_id_condition,
&count);
@@ -314,7 +348,7 @@ __blocking int asg_i_reset (
}
else
{
- *reset_count = count;
+ *records_reset = count;
rc = ASG_SUCCESS;
}
triton_error_destroy(tret);
diff --git a/code/src/asg/asg-internal.hae b/code/src/asg/asg-internal.hae
index cd4b616..f6a1ef5 100644
--- a/code/src/asg/asg-internal.hae
+++ b/code/src/asg/asg-internal.hae
@@ -46,22 +46,32 @@ __blocking int asg_i_write (
const void * data,
asg_size_t * transferred);
-__blocking int asg_i_read (
- asg_instance_t instance,
- asg_location_t location,
- asg_container_id_t container,
- asg_object_id_t object,
- asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
- asg_flags_t flags,
- asg_update_id_t version_condition,
- void * buf,
- size_t bufsize,
- asg_size_t * transferred,
- asg_record_info_t * record_buf,
- asg_size_t * record_buf_transferred,
- asg_update_id_t * version_info);
+__blocking int asg_i_read_sequence (
+ asg_instance_t instance,
+ asg_location_t location,
+ asg_container_id_t container,
+ asg_object_id_t object,
+ asg_fork_id_t fork,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
+ asg_size_t record_expected_len,
+ asg_flags_t flags,
+ asg_update_id_t update_id_condition,
+ const void * buf,
+ asg_record_info_t * record_info);
+
+__blocking int asg_i_read_one (
+ asg_instance_t instance,
+ asg_location_t location,
+ asg_container_id_t container,
+ asg_object_id_t object,
+ asg_fork_id_t fork,
+ asg_record_id_t record,
+ asg_flags_t flags,
+ asg_update_id_t update_id_condition,
+ void * buf,
+ size_t buf_size,
+ asg_record_info_t * record_info);
/**
* Punched is exactly like write but writes zero length records.
@@ -87,12 +97,11 @@ __blocking int asg_i_reset (
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
- asg_size_t recordlen,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
asg_flags_t flags,
asg_update_id_t update_id_condition,
- asg_size_t * reset_count);
+ asg_size_t * records_reset);
__blocking int asg_i_probe_system (
asg_instance_t instance,
diff --git a/code/src/asg/asg.c b/code/src/asg/asg.c
index edc2688..5cdccd1 100644
--- a/code/src/asg/asg.c
+++ b/code/src/asg/asg.c
@@ -1,5 +1,6 @@
#include <aesop/aesop.h>
#include <src/common/triton-error.h>
+#include <src/common/triton-debug.h>
#include <include/asg.h>
#include <src/asg/asg-internal.h>
@@ -33,6 +34,46 @@ int asg_finalize (asg_instance_t instance)
ASG_HANDLE_AE_DEFAULT(asg_i_finalize, instance)
}
+int asg_read_one (
+ asg_instance_t instance,
+ asg_location_t location,
+ asg_container_id_t container,
+ asg_object_id_t object,
+ asg_fork_id_t fork,
+ asg_record_id_t record,
+ asg_flags_t flags,
+ asg_update_id_t update_id_condition,
+ void * buf,
+ size_t buf_size,
+ asg_record_info_t * record_info)
+{
+ ASG_HANDLE_AE_DEFAULT(asg_i_read_one, instance, location, container,
+ object, fork, record, flags, update_id_condition, buf, buf_size,
+ record_info);
+}
+
+int asg_read_sequence (
+ asg_instance_t instance,
+ asg_location_t location,
+ asg_container_id_t container,
+ asg_object_id_t object,
+ asg_fork_id_t fork,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
+ asg_size_t record_expected_length,
+ asg_flags_t flags,
+ asg_update_id_t update_id_condition,
+ void * buf,
+ asg_record_info_t * record_info)
+{
+ ASG_HANDLE_AE_DEFAULT(asg_i_read_sequence, instance, location,
+ container, object,
+ fork, record_start, record_count, record_expected_length, flags,
+ update_id_condition, buf, record_info)
+}
+
+
+#if 0
int asg_read (
asg_instance_t instance,
asg_location_t location,
@@ -55,6 +96,7 @@ int asg_read (
bufsize, transferred, record_buf, record_buf_transferred,
version_info)
}
+#endif
int asg_write (
asg_instance_t instance,
@@ -100,16 +142,15 @@ int asg_reset (
asg_container_id_t container,
asg_object_id_t object,
asg_fork_id_t fork,
- asg_record_id_t start_record,
- asg_size_t recordcount,
- asg_size_t recordlen,
+ asg_record_id_t record_start,
+ asg_size_t record_count,
asg_flags_t flags,
asg_update_id_t update_id_condition,
- asg_size_t * reset_count)
+ asg_size_t * records_reset)
{
ASG_HANDLE_AE_DEFAULT(asg_i_reset, instance, location, container, object,
- fork, start_record, recordcount, recordlen, flags,
- update_id_condition, reset_count)
+ fork, record_start, record_count, flags,
+ update_id_condition, records_reset)
}
int asg_probe_system (
diff --git a/code/src/common/Makefile.subdir b/code/src/common/Makefile.subdir
index f7a406f..66453d4 100644
--- a/code/src/common/Makefile.subdir
+++ b/code/src/common/Makefile.subdir
@@ -24,7 +24,8 @@ BUILT_SOURCES += \
src/common/triton-error-decls.h \
src/common/triton-error-defs.c \
src/common/triton-debug-decls.h \
- src/common/triton-debug-defs.c
+ src/common/triton-debug-defs.c \
+ src/common/resources/signal/signal.h
AE_SRC += \
src/common/resources/signal/signal.ae \
diff --git a/code/src/common/triton-error.spec b/code/src/common/triton-error.spec
index 56d5cfe..a2ccca2 100644
--- a/code/src/common/triton-error.spec
+++ b/code/src/common/triton-error.spec
@@ -67,3 +67,4 @@ ERR_SSM_UNKNOWN "Unknown SSM error"
ERR_SSM_TRSPT_TYPE_UNAVAILABLE "SSM transport type unavailable at this endpoint"
ERR_NAME_HASH_COLLISION "Name hash collision"
ERR_NET_MODULE_SHUTTING_DOWN "Operation forbidden because net layer is shutting down"
+ERR_RECORDLEN_INVAL "Record length invalid"
diff --git a/code/src/replicated-osd/rosd-create.ae b/code/src/replicated-osd/rosd-create.ae
index d350c28..af53f79 100644
--- a/code/src/replicated-osd/rosd-create.ae
+++ b/code/src/replicated-osd/rosd-create.ae
@@ -61,6 +61,16 @@ static __blocking triton_ret_t triton_rpc_rosd_create(hg_handle_t handle)
"replication_factor: %d, niid: %llu\n",
in.container, in.oid, in.flags, in.replication_factor, llu(in.niid));
+ /* argument checking */
+ if(in.container == 0 || in.oid == 0)
+ {
+ out.tret = TRITON_ERR_INVAL;
+ triton_error_msg("Error: cannot create object in container or object 0.\n");
+ triton_mercury_respond(handle, &out);
+ HG_Destroy(handle);
+ return(TRITON_SUCCESS);
+ }
+
out.tret = TRITON_SUCCESS;
/* replication_factor of 0 -> defer to system default */
@@ -94,8 +104,11 @@ static __blocking triton_ret_t triton_rpc_rosd_create(hg_handle_t handle)
next_addr = addr_array[my_position+1];
}
- out.tret = rosd_create_do_work(in.container, in.oid, in.flags,
- actual_rep_factor, my_position, next_addr, in.niid, 1);
+ if(!triton_is_error(out.tret))
+ {
+ out.tret = rosd_create_do_work(in.container, in.oid, in.flags,
+ actual_rep_factor, my_position, next_addr, in.niid, 1);
+ }
triton_mercury_respond(handle, &out);
diff --git a/code/src/replicated-osd/rosd-read.ae b/code/src/replicated-osd/rosd-read.ae
index 43c6ec1..e174d3c 100644
--- a/code/src/replicated-osd/rosd-read.ae
+++ b/code/src/replicated-osd/rosd-read.ae
@@ -5,6 +5,7 @@
#include <triton-hash.h>
#include <mercury_macros.h>
#include <mercury_proc.h>
+#include <hoss.hae>
#include "src/replicated-osd/rosd.hae"
#include "src/replicated-osd/rosd-internal.hae"
@@ -18,57 +19,116 @@
#include "src/remote/mercury-encode.h"
#include "src/system-state/system-state.hae"
-/* this is the number of concurrent buffers/transfers that the ROSD will
- * keep in flight for a single read operation
- */
+/* see rosd.ae for management of read pipeline parameters */
extern int64_t rosd_read_xfer_pipeline_depth;
+extern int64_t rosd_read_buffer_size;
+extern struct buffer_mgmt_instance* rosd_read_buffers;
+extern struct hoss_ctx* g_hoss_ctx;
+
/* Mercury RPC structures for rosd_read */
-MERCURY_GEN_PROC(triton_rpc_rosd_read_in_t,
+MERCURY_GEN_PROC(triton_rpc_rosd_read_sequence_in_t,
((uint64_t)(container))\
((uint64_t)(object))\
((uint64_t)(fork))\
- ((uint64_t)(start_record))\
- ((uint64_t)(recordcount))\
+ ((uint64_t)(record_start))\
+ ((uint64_t)(record_count))\
+ ((uint64_t)(record_expected_len))\
((uint32_t)(flags))\
- ((hg_bulk_t)(bulk_handle))\
- ((uint64_t)(buf_size)))
+ ((uint64_t)(update_id_condition))\
+ ((hg_bulk_t)(bulk_handle)))
-MERCURY_GEN_PROC(triton_rpc_rosd_read_out_t,
- ((triton_ret_t)(tret))\
- ((uint64_t)(data_size))\
- ((uint64_t)(num_rec_infos)))
+MERCURY_GEN_PROC(triton_rpc_rosd_read_one_in_t,
+ ((uint64_t)(container))\
+ ((uint64_t)(object))\
+ ((uint64_t)(fork))\
+ ((uint64_t)(record))\
+ ((uint32_t)(flags))\
+ ((uint64_t)(update_id_condition))\
+ ((hg_bulk_t)(bulk_handle)))
-static hg_id_t rpc_rosd_read_id;
-static __blocking triton_ret_t triton_rpc_rosd_read(hg_handle_t handle);
-static hg_return_t triton_rpc_rosd_read_handler(hg_handle_t handle);
-extern struct buffer_mgmt_instance* rosd_read_buffers;
+/* TODO: record_info fields are split into separate arguments here for
+ * encoding simplicity, but it might be cleaner to provide an encoder for
+ * them in struct form.
+ */
+MERCURY_GEN_PROC(triton_rpc_rosd_read_sequence_out_t,
+ ((triton_ret_t)(tret))\
+ ((uint64_t)(record_id))\
+ ((uint64_t)(num_records))\
+ ((uint64_t)(record_len))\
+ ((uint64_t)(record_update_id)))
-__blocking triton_ret_t remote_triton_rpc_rosd_read(
+MERCURY_GEN_PROC(triton_rpc_rosd_read_one_out_t,
+ ((triton_ret_t)(tret))\
+ ((uint64_t)(record_len))\
+ ((uint64_t)(record_update_id)))
+
+static hg_id_t rpc_rosd_read_sequence_id, rpc_rosd_read_one_id;
+static __blocking triton_ret_t triton_rpc_rosd_read_sequence(hg_handle_t handle);
+static __blocking triton_ret_t triton_rpc_rosd_read_one(hg_handle_t handle);
+static hg_return_t triton_rpc_rosd_read_sequence_handler(hg_handle_t handle);
+static hg_return_t triton_rpc_rosd_read_one_handler(hg_handle_t handle);
+static __blocking triton_ret_t rosd_read_run_pipeline(
+ hg_handle_t handle,
+ na_addr_t src_addr,
+ triton_rpc_rosd_read_sequence_in_t* in,
+ uint64_t *records_read,
+ uint64_t *new_update_id);
+static __blocking triton_ret_t rosd_handle_read_segment(
+ hg_handle_t handle,
+ na_addr_t src_addr,
+ triton_rpc_rosd_read_sequence_in_t* in,
+ uint64_t start_record,
+ int64_t remote_offset,
+ uint64_t recordcount,
+ uint64_t *records_read,
+ uint64_t *new_update_id);
+static __blocking triton_ret_t rosd_read_one_local_storage(
+ uint64_t container,
+ uint64_t object,
+ uint64_t fork,
+ uint64_t record,
+ uint32_t flags,
+ uint64_t update_id_condition,
+ void * buf,
+ uint64_t buf_sz,
+ uint64_t *record_len,
+ uint64_t *new_update_id);
+static __blocking triton_ret_t rosd_read_sequence_local_storage(
+ uint64_t container,
+ uint64_t object,
+ uint64_t fork,
+ uint64_t record_start,
+ uint64_t record_count,
+ uint64_t record_expected_len,
+ char* buffer,
+ uint32_t flags,
+ uint64_t update_id_condition,
+ uint64_t *records_read,
+ uint64_t *new_update_id);
+
+__blocking triton_ret_t remote_triton_rpc_rosd_read_sequence(
uint64_t container,
uint64_t object,
uint64_t fork,
- uint64_t start_record,
- uint64_t recordcount,
+ uint64_t record_start,
+ uint64_t record_count,
+ uint64_t record_expected_len,
int flags,
+ uint64_t update_id_condition,
void * buf,
- uint64_t buf_size,
- uint64_t * version_info,
- uint64_t * transferred,
- struct rosd_record_info * record_info,
- uint64_t * record_info_xferred)
+ struct rosd_record_info* record_info)
{
- triton_rpc_rosd_read_out_t out;
- triton_rpc_rosd_read_in_t in;
+ triton_rpc_rosd_read_sequence_out_t out;
+ triton_rpc_rosd_read_sequence_in_t in;
int ret;
triton_ret_t tret;
na_addr_t addr;
int position;
- hg_handle_t handle;
+ hg_handle_t handle = HG_HANDLE_NULL;
+ size_t bulk_size = record_count*record_expected_len;
struct hg_info *info;
- void * bulk_bufs[2];
- hg_size_t bulk_sizes[2];
/* TODO: ability read from replicas as well */
/* we contact the first server (master) for oid */
@@ -81,19 +141,15 @@ __blocking triton_ret_t remote_triton_rpc_rosd_read(
in.container = container;
in.object = object;
in.fork = fork;
- in.start_record = start_record;
- in.recordcount = recordcount;
- in.flags = flags;
- in.buf_size = buf_size;
+ in.record_start = record_start;
+ in.record_count = record_count;
+ in.record_expected_len = record_expected_len;
in.flags = flags;
+ in.update_id_condition = update_id_condition;
in.bulk_handle = HG_BULK_NULL;
- bulk_bufs[0] = buf;
- bulk_bufs[1] = record_info;
- bulk_sizes[0] = buf_size;
- bulk_sizes[1] = recordcount * sizeof(*record_info);
- triton_mercury_run_rpc_with_bulk(addr, rpc_rosd_read_id, handle,
- &in, &out, in.bulk_handle, 2, bulk_bufs, bulk_sizes, 0,
+ triton_mercury_run_rpc_with_bulk(addr, rpc_rosd_read_sequence_id, handle,
+ &in, &out, in.bulk_handle, 1, &buf, &bulk_size, 0,
HG_BULK_WRITE_ONLY);
triton_mercury_addr_free(addr);
@@ -101,13 +157,10 @@ __blocking triton_ret_t remote_triton_rpc_rosd_read(
tret = triton_error_dup(out.tret);
if(!triton_is_error(tret))
{
- *transferred = out.data_size;
- *record_info_xferred = out.num_rec_infos;
- }
- else
- {
- *transferred = 0;
- *version_info = 0;
+ record_info->record_id = out.record_id;
+ record_info->num_records = out.num_records;
+ record_info->record_len = out.record_len;
+ record_info->record_update_id = out.record_update_id;
}
HG_Free_output(handle, &out);
@@ -118,88 +171,425 @@ __blocking triton_ret_t remote_triton_rpc_rosd_read(
}
-#if 0
-/* TODO: convert me to rosd */
-static __blocking triton_ret_t rosd_read_xfer_one_buffer(
- struct hg_info *hg_info,
- char* tmp_buffer,
- uint128_t oid,
+__blocking triton_ret_t remote_triton_rpc_rosd_read_one(
+ uint64_t container,
+ uint64_t object,
+ uint64_t fork,
+ uint64_t record,
+ uint32_t flags,
+ uint64_t update_id_condition,
+ void * buf,
+ size_t buf_size,
+ uint64_t record_len,
+ uint64_t * record_update_id) {
+
+ triton_rpc_rosd_read_one_in_t in;
+ triton_rpc_rosd_read_one_out_t out;
+
+ hg_handle_t handle = HG_HANDLE_NULL;
+ na_addr_t addr;
+ int position;
+
+ triton_ret_t tret;
+
+ tret = triton_oid_to_addrs(object, 1, &position, &addr);
+ if(triton_is_error(tret))
+ return(tret);
+
+ in.container = container;
+ in.object = object;
+ in.fork = fork;
+ in.record = record;
+ in.flags = flags;
+ in.update_id_condition = update_id_condition;
+ in.bulk_handle = HG_BULK_NULL;
+
+ triton_mercury_run_rpc_with_bulk(addr, rpc_rosd_read_one_id, handle, &in,
+ &out, in.bulk_handle, 1, &buf, &buf_size, 0, HG_BULK_WRITE_ONLY);
+
+ tret = triton_error_dup(out.tret);
+
+ triton_mercury_addr_free(addr);
+ HG_Free_output(handle, &out);
+ HG_Destroy(handle);
+
+ return tret;
+}
+
+/* TODO: merge this with other functions */
+
+static __blocking triton_ret_t rosd_read_one_local_storage(
+ uint64_t container,
+ uint64_t object,
+ uint64_t fork,
+ uint64_t record,
+ uint32_t flags,
+ uint64_t update_id_condition,
+ void * buf,
+ uint64_t buf_sz,
+ uint64_t *record_len,
+ uint64_t *new_update_id){
+
+ triton_ret_t tret;
+
+ struct hoss_grp *grp;
+ struct hoss_oid soid;
+ hoss_eid_t soid_eids[3];
+ hoss_size_t nbytes;
+ int ret, hoss_rc;
+ hoss_size_t nrecs;
+ hoss_update_t update_info;
+ struct hoss_record_info rec_info;
+
+ soid.nids = 3;
+ soid.ids = soid_eids;
+ soid_eids[0] = container;
+ soid_eids[1] = object;
+ soid_eids[2] = fork;
+
+ /* TODO: handle canceled error code (or whatever hoss will report to
+ * indicate that a txn failed and should be retried)
+ */
+ ret = hoss_begin(&grp, NULL, HOSS_NONE, g_hoss_ctx);
+ if(ret != 0)
+ return(TRITON_ERR_IO);
+
+ ret = hoss_read(&soid, record, 1, HOSS_NONE, update_id_condition, buf,
+ buf_sz, &rec_info, 1, grp, &nbytes, &nrecs, &update_info,
+ &hoss_rc);
+
+ if (ret != 0)
+ tret = TRITON_ERR_IO;
+ /* check for short buffer */
+ else if (hoss_rc == ENOSPC) {
+ tret = TRITON_ERR_RECV_TOO_SMALL;
+ }
+ /* check for other io errors */
+ else if (hoss_rc != 0)
+ tret = TRITON_ERR_IO;
+ /* check for nonexistent record */
+ else if (nrecs == 0)
+ tret = TRITON_ERR_NOENT;
+ else
+ tret = TRITON_SUCCESS;
+
+ hoss_end(grp, 1);
+ return tret;
+}
+
+static __blocking triton_ret_t rosd_read_sequence_local_storage(
+ uint64_t container,
+ uint64_t object,
uint64_t fork,
- int64_t size,
- int64_t offset,
+ uint64_t record_start,
+ uint64_t record_count,
+ uint64_t record_expected_len,
+ char* buffer,
+ uint32_t flags,
+ uint64_t update_id_condition,
+ uint64_t *records_read,
+ uint64_t *new_update_id)
+{
+ struct hoss_grp *grp;
+ struct hoss_oid soid;
+ hoss_eid_t soid_eids[3];
+ hoss_size_t nbytes;
+ int ret, hoss_rc;
+ hoss_size_t n_recs;
+ hoss_update_t update_info;
+
+ /* NOTE: we provide two hoss_record_info structs here. In normal
+ * operation we only expect 1 of them to be filled in. The second one
+ * is there to distinguish between cases where we hit a short read (or a
+ * hole) vs. cases where we encountered a record of the wrong size mixed
+ * into the sequence.
+ */
+ struct hoss_record_info rec_array[2];
+
+ /* TODO: handle canceled error code (or whatever hoss will report to
+ * indicate that a txn failed and should be retried)
+ */
+ ret = hoss_begin(&grp, NULL, HOSS_NONE, g_hoss_ctx);
+ if(ret != 0)
+ {
+ return(TRITON_ERR_IO);
+ }
+
+ soid.nids = 3;
+ soid.ids = soid_eids;
+ soid_eids[0] = container;
+ soid_eids[1] = object;
+ soid_eids[2] = fork;
+
+ ret = hoss_read(&soid, record_start, record_count, HOSS_NONE, update_id_condition, buffer, record_count*record_expected_len, rec_array, 2, grp, &nbytes, &n_recs, &update_info, &hoss_rc);
+
+ /* check for fundamental I/O error */
+ if(ret != 0 || hoss_rc != 0)
+ {
+ triton_error_msg("hoss_read() failure, return code %d\n", ret);
+ hoss_end(grp, 0);
+ return(TRITON_ERR_IO);
+ }
+ /* check for unexpected record size */
+ if(n_recs == 2 || (n_recs ==1 && rec_array[0].reclen != record_expected_len))
+ {
+ /* there was either a mix of record sizes, or one record size that didn't
+ * match expectations
+ */
+ triton_error_msg("hoss_read() found unexpected record size.\n");
+ hoss_end(grp, 0);
+ return(TRITON_ERR_RECORDLEN_INVAL);
+ }
+
+ ret = hoss_end(grp, 1);
+ if(ret != 0)
+ {
+ return(TRITON_ERR_IO);
+ }
+
+ /* report amount read */
+ if(nbytes)
+ {
+ assert(n_recs == 1);
+ *records_read = rec_array[0].nrecs;
+ }
+ else
+ {
+ *records_read = 0;
+ }
+
+ return(TRITON_SUCCESS);
+}
+
+
+static __blocking triton_ret_t rosd_handle_read_segment(
+ hg_handle_t handle,
+ na_addr_t src_addr,
+ triton_rpc_rosd_read_sequence_in_t* in,
+ uint64_t start_record,
int64_t remote_offset,
- hg_bulk_t remote_bulk_handle,
- int64_t *out_size,
- uint64_t *txn_number)
+ uint64_t recordcount,
+ uint64_t *records_read,
+ uint64_t *new_update_id)
{
+ char *buffer;
triton_ret_t tret;
+ int64_t outsize;
+ struct buffer_mgmt_token *token;
hg_bulk_t bulk_handle = HG_BULK_NULL;
int ret;
+ struct hg_info *info = NULL;
+ size_t bulk_size;
- tret = tosd_read_versioned(
- oid,
- fork,
- &tmp_buffer,
- &size,
- 1,
- &offset,
- &size,
- 1,
- out_size,
- 0,
- txn_number);
-
- if(triton_is_error(tret))
+ tret = buffer_mgmt_alloc(rosd_read_buffers, recordcount*in->record_expected_len, &outsize, &buffer, &token);
+ assert(outsize == recordcount*in->record_expected_len);
+ if(tret != TRITON_SUCCESS)
{
return(tret);
}
- /* zero byte read; don't try to do bulk transfer */
- if((*out_size) == 0)
- {
- return(TRITON_SUCCESS);
- }
+ info = HG_Get_info(handle);
+ bulk_size = outsize;
- /* TODO: fix types for arguments here... */
- ret = HG_Bulk_create(hg_info->hg_bulk_class, 1, (void**)&tmp_buffer, (size_t*)out_size, HG_BULK_READ_ONLY, &bulk_handle);
+ ret = HG_Bulk_create(info->hg_bulk_class, 1, (void**)&buffer,
+ &bulk_size, HG_BULK_READ_ONLY, &bulk_handle);
if(ret != HG_SUCCESS)
{
triton_error_msg("HG_Bulk_create() failure.\n");
+ buffer_mgmt_free(token);
return(TRITON_ERR_NOMEM);
}
- ret = hg_rsc_bulk_transfer(hg_info->bulk_context, HG_BULK_PUSH,
- hg_info->addr, remote_bulk_handle, remote_offset, bulk_handle, 0, *out_size);
- if(ret != HG_SUCCESS)
+ /* NOTE: the start record and count is dictated by function call arguments
+ * (per segment) rather than by in request (entire operation)
+ */
+ tret = rosd_read_sequence_local_storage(in->container, in->object,
+ in->fork, start_record, recordcount, in->record_expected_len,
+ buffer, in->flags, in->update_id_condition, records_read,
+ new_update_id);
+ if(triton_is_error(tret))
{
- triton_error_msg("HG_Bulk_write() failure.\n");
+ triton_error_msg("rosd_read_do_work() failure.\n");
HG_Bulk_free(bulk_handle);
- return(TRITON_ERR_UNKNOWN);
+ buffer_mgmt_free(token);
+ return(tret);
}
+ assert(*records_read <= recordcount);
+
+ /* only bulk xfer actual bytes that were read from storage */
+ bulk_size = (*records_read) * in->record_expected_len;
+ if(bulk_size > 0)
+ {
+ ret = hg_rsc_bulk_transfer(info->bulk_context, HG_BULK_PUSH, src_addr,
+ in->bulk_handle, remote_offset, bulk_handle, 0, bulk_size);
+ }
+ else
+ {
+ ret = HG_SUCCESS;
+ }
+ /* TODO: properly handle cancellations */
+ assert(ret != AE_ERR_CANCELLED);
+ if (ret != HG_SUCCESS){
+ triton_error_msg("hg_rsc_bulk_transfer() failure.\n");
+ HG_Bulk_free(bulk_handle);
+ buffer_mgmt_free(token);
+ return(TRITON_ERR_UNKNOWN);
+ }
HG_Bulk_free(bulk_handle);
+ buffer_mgmt_free(token);
+
return(tret);
}
-#endif
-static __blocking triton_ret_t triton_rpc_rosd_read(hg_handle_t handle)
+static __blocking triton_ret_t rosd_read_run_pipeline(
+ hg_handle_t handle,
+ na_addr_t src_addr,
+ triton_rpc_rosd_read_sequence_in_t* in,
+ uint64_t *records_read,
+ uint64_t *new_update_id)
+{
+ int64_t current_remote_offset = 0;
+ uint64_t current_record = in->record_start;
+ triton_ret_t out_tret = TRITON_SUCCESS;
+ triton_mutex_t pipeline_mutex;
+ uint64_t records_per_buffer;
+ int got_short = 0;
+
+ triton_mutex_init(&pipeline_mutex, NULL);
+
+ records_per_buffer = rosd_read_buffer_size/in->record_expected_len;
+
+ if(records_per_buffer == 0)
+ {
+ triton_error_msg("record length of %llu is too large.\n", llu(in->record_expected_len));
+ triton_mutex_destroy(&pipeline_mutex);
+ return(TRITON_ERR_INVAL);
+ }
+
+ /* NOTE: we start out by assuming that all requested records will be read,
+ * and we reduce the count if portions of the pipeline come up short
+ */
+ *records_read = in->record_count;
+ *new_update_id = 0;
+
+ pwait
+ {
+ pprivate triton_ret_t b_tret;
+ pprivate int i;
+ pprivate uint64_t this_start_record;
+ pprivate int64_t this_remote_offset;
+ pprivate uint64_t this_recordcount;
+ pprivate uint64_t this_records_read;
+ pprivate uint64_t this_new_update_id;
+
+ /* One pbranch for pipeline depth. Each pbranch can be though of
+ * like a worker in a thread pool model.
+ */
+ for(i=0; i<rosd_read_xfer_pipeline_depth; i++)
+ {
+ pbranch
+ {
+ /* Each pbranch keeps greedily looking for the "next"
+ * segment to transfer. We don't pre-assign pipeline
+ * segments to each pbranch, because some segments might run
+ * slower than others and lead to stragglers. Bulk
+ * transfers in Mercury can be performed in any order; we
+ * just make a best effort to issue HOSS operations
+ * sequentially.
+ */
+ while(1)
+ {
+ /* calculate which segment this iteration will work on */
+ triton_mutex_lock(&pipeline_mutex);
+ /* If we reach the end of the pipeline or someone has
+ * encountered an I/O error, or we already know it is a short
+ * read, then stop here.
+ */
+ if(current_record >= (in->record_start + in->record_count)
+ || triton_is_error(out_tret) || got_short)
+ {
+ triton_mutex_unlock(&pipeline_mutex);
+ pbreak;
+ }
+
+ this_start_record = current_record;
+ current_record += records_per_buffer;
+ this_remote_offset = current_remote_offset;
+ current_remote_offset += records_per_buffer * in->record_expected_len;
+ this_recordcount = records_per_buffer;
+ if(this_recordcount + this_start_record > in->record_start + in->record_count)
+ this_recordcount = (in->record_start + in->record_count) - this_start_record;
+ triton_mutex_unlock(&pipeline_mutex);
+
+ triton_debug(triton_dbg_rosd, "rosd_read_run_pipeline() pbranch %d working on records %llu to %llu, remote offset %lld\n", i, llu(this_start_record), llu(this_start_record+this_recordcount), lld(this_remote_offset));
+
+ b_tret = rosd_handle_read_segment(
+ handle,
+ src_addr,
+ in,
+ this_start_record,
+ this_remote_offset,
+ this_recordcount,
+ &this_records_read,
+ &this_new_update_id);
+
+ triton_mutex_lock(&pipeline_mutex);
+ if(triton_is_error(b_tret))
+ {
+ out_tret = b_tret;
+ }
+ else
+ {
+ /* was this specific segment short? */
+ if(this_records_read < this_recordcount)
+ {
+ got_short = 1;
+ if((*records_read) >
+ ((this_start_record - in->record_start) + this_records_read))
+ {
+ /* this is the lowest short read encountered so
+ * far
+ */
+ *records_read = this_start_record -
+ in->record_start + this_records_read;
+ }
+ }
+ if(this_records_read)
+ {
+ /* if we read something, consider updating aggregate
+ * update id
+ */
+ if(*new_update_id == 0)
+ *new_update_id = this_new_update_id;
+ else if(*new_update_id != this_new_update_id)
+ *new_update_id = ROSD_UPDATE_ID_MIXED;
+ }
+ }
+ triton_mutex_unlock(&pipeline_mutex);
+ }
+ }
+ }
+ }
+
+ triton_mutex_destroy(&pipeline_mutex);
+ return(out_tret);
+}
+
+static __blocking triton_ret_t triton_rpc_rosd_read_sequence(hg_handle_t handle)
{
- triton_rpc_rosd_read_out_t out;
- triton_rpc_rosd_read_in_t in;
+ triton_rpc_rosd_read_sequence_out_t out;
+ triton_rpc_rosd_read_sequence_in_t in;
+ na_addr_t src_addr;
na_addr_t next_addr;
int my_position;
- int ret;
- int64_t size_remaining = 0;
- int64_t remote_offset = 0;
- int64_t local_offset = 0;
- aesop_sem_t sem;
- triton_mutex_t output_mutex;
struct hg_info *hg_info = NULL;
triton_debug(triton_dbg_rosd, "Called triton_rpc_rosd_read()\n");
hg_info = HG_Get_info(handle);
+ src_addr = hg_info->addr;
triton_mercury_get_input(handle, &in, &out);
@@ -223,138 +613,104 @@ static __blocking triton_ret_t triton_rpc_rosd_read(hg_handle_t handle)
return(TRITON_SUCCESS);
}
- out.tret = TRITON_ERR_NOT_IMPLEMENTED;
- out.data_size = 0;
- out.num_rec_infos = 0;
+ out.tret = rosd_read_run_pipeline(handle, src_addr, &in, &out.num_records,
+ &out.record_update_id);
+ /* NOTE: assume other output fields ignored if error code set */
+ out.record_len = in.record_expected_len;
+ out.record_id = in.record_start;
+
triton_mercury_respond(handle, &out);
HG_Destroy(handle);
return(TRITON_SUCCESS);
+}
+TRITON_DEFINE_RPC_HANDLER(triton_rpc_rosd_read_sequence)
-#if 0
- /* TODO: convert to rosd */
- size_remaining = in.size;
- local_offset = in.offset;
- remote_offset = 0;
- aesop_sem_init(&sem, 1);
- triton_mutex_init(&output_mutex, NULL);
+/* placeholder */
+static __blocking triton_ret_t triton_rpc_rosd_read_one(hg_handle_t handle)
+{
+ /* input/outputs */
+ triton_rpc_rosd_read_one_out_t out;
+ triton_rpc_rosd_read_one_in_t in;
- pwait
- {
- pprivate int i = 0;
- pprivate int64_t this_size = 0;
- pprivate int64_t this_out_size = 0;
- pprivate int64_t this_local_offset = 0;
- pprivate int64_t this_remote_offset = 0;
- pprivate struct buffer_mgmt_token *token = NULL;
- pprivate char* buffer = NULL;
- pprivate triton_ret_t tret;
- pprivate uint64_t this_txn_number = 0;
-
- //triton_debug(triton_dbg_rosd, "triton_rpc_rosd_read() in pwait, size_remaining: %d.\n", size_remaining);
- for(i=0; i<rosd_read_xfer_pipeline_depth; i++)
- {
- pbranch
- {
- do
- {
- /* calculate how much to transfer in this step */
- /* note that this function protects shared variables
- * using a semaphore
- */
- //triton_debug(triton_dbg_rosd, "triton_rpc_rosd_read() pbranch %d calling setup_one_buffer().\n", i);
- tret = rosd_pipeline_setup_one_buffer(
- rosd_read_buffers,
- &sem,
- &size_remaining,
- &remote_offset,
- &local_offset,
- &this_size,
- &this_local_offset,
- &this_remote_offset,
- &token,
- &buffer);
- triton_mutex_lock(&output_mutex);
- if(triton_is_error(tret) && !triton_is_error(out.tret))
- {
- out.tret = tret;
- this_size = 0;
- }
- else if(triton_is_error(tret))
- {
- triton_error_destroy(tret);
- }
- triton_mutex_unlock(&output_mutex);
-
- //triton_debug(triton_dbg_rosd, "triton_rpc_rosd_read() pbranch %d finished setup_one_buffer(), size_remaining: %d.\n", i, size_remaining);
+ /* mercury */
+ na_addr_t next_addr;
+ int my_position;
+ struct hg_info *info = NULL;
+ hg_bulk_t bulk_handle = HG_BULK_NULL;
+ int ret;
- if(this_size > 0)
- {
- /* perform buffer transfer */
- //triton_debug(triton_dbg_rosd, "triton_rpc_rosd_read() pbranch %d calling xfer_one_buffer().\n", i);
- tret = rosd_read_xfer_one_buffer(
- hg_info,
- buffer,
- in.oid,
- in.oid_fork,
- this_size,
- this_local_offset,
- this_remote_offset,
- in.bulk_handle,
- &this_out_size,
- &this_txn_number);
-
- //triton_debug(triton_dbg_rosd, "triton_rpc_rosd_read() pbranch %d finished xfer_one_buffer().\n", i);
- /* accumulate results in rpc response */
- triton_mutex_lock(&output_mutex);
- if(triton_is_error(tret) && !triton_is_error(out.tret))
- {
- out.tret = tret;
- }
- else if(triton_is_error(tret))
- {
- triton_error_destroy(tret);
- }
- else
- {
- out.out_size += this_out_size;
- if(out.txn_number == 0 && out.txn_number == this_txn_number)
- out.txn_number = this_txn_number;
- else
- {
- /* TODO: what are we supposed to do on mixed txn
- * numbers?
- */
- out.txn_number = 0;
- }
- }
- triton_mutex_unlock(&output_mutex);
- }
- else
- {
- this_out_size = this_size;
- }
- }while(this_out_size > 0 && !triton_is_error(out.tret));
+ /* read buffer */
+ void * buffer;
+ struct buffer_mgmt_token *token = NULL;
+ int64_t insize;
+ int64_t outsize;
+ size_t bulk_size;
- if(token != NULL)
- {
- buffer_mgmt_free(token);
- }
- //triton_debug(triton_dbg_rosd, "triton_rpc_rosd_read() done with pbranch %d.\n", i);
- }
- }
+ triton_debug(triton_dbg_rosd, "Called triton_rpc_rosd_read_one()\n");
+
+ info = HG_Get_info(handle);
+
+ triton_mercury_get_input(handle, &in, &out);
+
+ /* TODO: ability to read from replicas */
+ /* check our position assuming a replication factor of 1 to make
+ * sure that we are the master for this object
+ */
+ out.tret = triton_oid_to_addrs(in.object, 1, &my_position, &next_addr);
+ if(triton_is_error(out.tret))
+ goto respond;
+ if (my_position != 0) {
+ out.tret = TRITON_ERR_WRONG_SERVER;
+ goto respond;
+ }
+
+ /* TODO: define a max record size so users can't blow us up with large max
+ * buffer sizes */
+ insize = HG_Bulk_get_size(in.bulk_handle);
+ out.tret = buffer_mgmt_alloc(rosd_read_buffers, insize, &outsize,
+ (char**)&buffer, &token);
+ /* TODO: what's the appropriate error if we request a record larger than
+ * the largest buffer */
+ if (triton_is_error(out.tret) || insize > outsize) {
+ out.tret = (triton_is_error(out.tret)) ? out.tret : TRITON_ERR_NOMEM;
+ goto respond;
}
- //triton_debug(triton_dbg_rosd, "triton_rpc_rosd_read() done with pwait.\n");
- aesop_sem_destroy(&sem);
- triton_mutex_destroy(&output_mutex);
+ /* TODO: is pipelining possible on a single record? doesn't look like it */
+ out.tret = rosd_read_one_local_storage(in.container, in.object, in.fork,
+ in.record, in.flags, in.update_id_condition, buffer,
+ (uint64_t)outsize, &out.record_len, &out.record_update_id);
+
+ if (out.tret == TRITON_SUCCESS) {
+ bulk_size = outsize;
+ ret = HG_Bulk_create(info->hg_bulk_class, 1, &buffer,
+ &bulk_size, HG_BULK_READ_ONLY, &bulk_handle);
+ if(ret != HG_SUCCESS) {
+ triton_error_msg("HG_Bulk_create() failure.\n");
+ out.tret = TRITON_ERR_NOMEM;
+ }
+ else {
+ ret = hg_rsc_bulk_transfer(info->bulk_context, HG_BULK_PUSH,
+ info->addr, in.bulk_handle, 0, bulk_handle, 0,
+ bulk_size);
+ assert(ret != AE_ERR_CANCELLED);
+ if (ret != HG_SUCCESS)
+ out.tret = TRITON_ERR_UNKNOWN;
+ }
+ }
+respond:
triton_mercury_respond(handle, &out);
+
+ if (token != NULL)
+ buffer_mgmt_free(token);
+ if (bulk_handle != HG_BULK_NULL)
+ HG_Bulk_free(bulk_handle);
HG_Destroy(handle);
- return(TRITON_SUCCESS);
-#endif
+ return TRITON_SUCCESS;
}
-TRITON_DEFINE_RPC_HANDLER(triton_rpc_rosd_read)
+TRITON_DEFINE_RPC_HANDLER(triton_rpc_rosd_read_one)
void triton_rpc_rosd_read_register(void)
{
@@ -362,10 +718,17 @@ void triton_rpc_rosd_read_register(void)
hg_class = triton_mercury_engine_get_class();
- rpc_rosd_read_id = MERCURY_REGISTER(hg_class, "triton_rpc_rosd_read",
- triton_rpc_rosd_read_in_t,
- triton_rpc_rosd_read_out_t,
- triton_rpc_rosd_read_handler);
+ rpc_rosd_read_sequence_id = MERCURY_REGISTER(hg_class, "triton_rpc_rosd_read_sequence",
+ triton_rpc_rosd_read_sequence_in_t,
+ triton_rpc_rosd_read_sequence_out_t,
+ triton_rpc_rosd_read_sequence_handler);
+
+ rpc_rosd_read_one_id = MERCURY_REGISTER(hg_class,
+ "triton_rpc_rosd_read_one",
+ triton_rpc_rosd_read_one_in_t,
+ triton_rpc_rosd_read_one_out_t,
+ triton_rpc_rosd_read_one_handler);
+
return;
}
diff --git a/code/src/replicated-osd/rosd-reset.ae b/code/src/replicated-osd/rosd-reset.ae
index f9d6913..21dfba0 100644
--- a/code/src/replicated-osd/rosd-reset.ae
+++ b/code/src/replicated-osd/rosd-reset.ae
@@ -19,7 +19,6 @@ __blocking triton_ret_t __remote_triton_rpc_rosd_reset(
uint64_t fork,
uint64_t record,
uint64_t rcount,
- uint64_t rlen,
int flags,
uint64_t condition,
uint64_t *reset_count,
@@ -31,7 +30,6 @@ MERCURY_GEN_PROC(triton_rpc_rosd_reset_in_t,
((uint64_t)(fork)) \
((uint64_t)(record)) \
((uint64_t)(rcount)) \
- ((uint64_t)(rlen)) \
((int64_t)(flags)) \
((uint64_t)(condition)) \
((int32_t)(int_pos)))
@@ -115,7 +113,6 @@ static __blocking triton_ret_t triton_rpc_rosd_reset(hg_handle_t handle)
in.fork,
in.record,
in.rcount,
- in.rlen,
in.flags,
in.condition,
&out.reset_count,
@@ -135,7 +132,6 @@ static __blocking triton_ret_t triton_rpc_rosd_reset(hg_handle_t handle)
in.fork,
in.record,
in.rcount,
- in.rlen,
in.flags,
in.condition,
&out.reset_count,
@@ -178,7 +174,6 @@ __blocking triton_ret_t remote_triton_rpc_rosd_reset(
uint64_t fork,
uint64_t record,
uint64_t rcount,
- uint64_t rlen,
int flags,
uint64_t condition,
uint64_t *reset_count)
@@ -201,7 +196,6 @@ __blocking triton_ret_t remote_triton_rpc_rosd_reset(
fork,
record,
rcount,
- rlen,
flags,
condition,
reset_count,
@@ -220,7 +214,6 @@ __blocking triton_ret_t __remote_triton_rpc_rosd_reset(
uint64_t fork,
uint64_t record,
uint64_t rcount,
- uint64_t rlen,
int flags,
uint64_t condition,
uint64_t *reset_count,
@@ -237,7 +230,6 @@ __blocking triton_ret_t __remote_triton_rpc_rosd_reset(
in.fork = fork;
in.record = record;
in.rcount = rcount;
- in.rlen = rlen;
in.flags = flags;
in.condition = condition;
in.int_pos = position;
diff --git a/code/src/replicated-osd/rosd-write.ae b/code/src/replicated-osd/rosd-write.ae
index 30de2d8..fb3c0f7 100644
--- a/code/src/replicated-osd/rosd-write.ae
+++ b/code/src/replicated-osd/rosd-write.ae
@@ -116,7 +116,7 @@ __blocking triton_ret_t remote_triton_rpc_rosd_write(
}
-static __blocking triton_ret_t rosd_handle_segment(
+static __blocking triton_ret_t rosd_handle_write_segment(
hg_handle_t handle,
na_addr_t src_addr,
na_addr_t next_addr,
@@ -197,7 +197,6 @@ static __blocking triton_ret_t rosd_write_run_pipeline(
records_per_buffer = rosd_write_buffer_size/in->recordlen;
- /* TODO: check for zero byte writes somewhere earlier */
if(records_per_buffer == 0)
{
triton_error_msg("record length of %llu is too large.\n", llu(in->recordlen));
@@ -213,45 +212,63 @@ static __blocking triton_ret_t rosd_write_run_pipeline(
pprivate int64_t this_remote_offset;
pprivate uint64_t this_recordcount;
+ /* One pbranch for pipeline depth. Each pbranch can be though of
+ * like a worker in a thread pool model.
+ */
for(i=0; i<rosd_write_xfer_pipeline_depth; i++)
{
pbranch
{
- /* calculate which segment this iteration will work on */
- triton_mutex_lock(&pipeline_mutex);
- if(current_record >= (in->start_record + in->recordcount))
+ /* Each pbranch keeps greedily looking for the "next"
+ * segment to transfer. We don't pre-assign pipeline
+ * segments to each pbranch, because some segments might run
+ * slower than others and lead to stragglers. Bulk
+ * transfers in Mercury can be performed in any order; we
+ * just make a best effort to issue HOSS operations
+ * sequentially.
+ */
+ while(1)
{
+ /* calculate which segment this iteration will work on */
+ triton_mutex_lock(&pipeline_mutex);
+ /* If we reach the end of the pipeline or someone has
+ * encountered an I/O error, then stop here.
+ */
+ if(current_record >= (in->start_record + in->recordcount)
+ || triton_is_error(out_tret))
+ {
+ triton_mutex_unlock(&pipeline_mutex);
+ pbreak;
+ }
+
+ this_start_record = current_record;
+ current_record += records_per_buffer;
+ this_remote_offset = current_remote_offset;
+ current_remote_offset += records_per_buffer * in->recordlen;
+ this_recordcount = records_per_buffer;
+ if(this_recordcount + this_start_record > in->start_record + in->recordcount)
+ this_recordcount = (in->start_record + in->recordcount) - this_start_record;
triton_mutex_unlock(&pipeline_mutex);
- pbreak;
- }
- this_start_record = current_record;
- current_record += records_per_buffer;
- this_remote_offset = current_remote_offset;
- current_remote_offset += records_per_buffer * in->recordlen;
- this_recordcount = records_per_buffer;
- if(this_recordcount + this_start_record > in->start_record + in->recordcount)
- this_recordcount = (in->start_record + in->recordcount) - this_start_record;
- triton_mutex_unlock(&pipeline_mutex);
-
- triton_debug(triton_dbg_rosd, "rosd_write_run_pipeline() pbranch %d working on records %llu to %llu, remote offset %lld\n", i, llu(this_start_record), llu(this_start_record+this_recordcount), lld(this_remote_offset));
-
- b_tret = rosd_handle_segment(
- handle,
- src_addr,
- next_addr,
- my_position,
- in,
- this_start_record,
- this_remote_offset,
- this_recordcount);
-
- triton_mutex_lock(&pipeline_mutex);
- if(triton_is_error(b_tret))
- {
- out_tret = b_tret;
+ triton_debug(triton_dbg_rosd, "rosd_write_run_pipeline() pbranch %d working on records %llu to %llu, remote offset %lld\n", i, llu(this_start_record), llu(this_start_record+this_recordcount), lld(this_remote_offset));
+
+ b_tret = rosd_handle_write_segment(
+ handle,
+ src_addr,
+ next_addr,
+ my_position,
+ in,
+ this_start_record,
+ this_remote_offset,
+ this_recordcount);
+
+ triton_mutex_lock(&pipeline_mutex);
+ if(triton_is_error(b_tret))
+ {
+ out_tret = b_tret;
+ }
+ triton_mutex_unlock(&pipeline_mutex);
}
- triton_mutex_unlock(&pipeline_mutex);
}
}
}
@@ -287,7 +304,7 @@ static __blocking triton_ret_t resolve_replication_factor(
* retried.
*/
- ret = hoss_begin(&grp, NULL, 0, g_hoss_ctx);
+ ret = hoss_begin(&grp, NULL, HOSS_MSYNC|HOSS_ORDERED, g_hoss_ctx);
if(ret != 0)
{
return(TRITON_ERR_IO);
@@ -299,7 +316,7 @@ static __blocking triton_ret_t resolve_replication_factor(
* factor here against what is stored on disk (if present)
*/
- ret = hoss_read(&soid, 0, 1, 0, 0, &ondisk_rep_factor, sizeof(ondisk_rep_factor), &ribuf, 1, grp, &nbytes, &nrec_infos, &update_info, &hoss_rc);
+ ret = hoss_read(&soid, 0, 1, HOSS_NONE, 0, &ondisk_rep_factor, sizeof(ondisk_rep_factor), &ribuf, 1, grp, &nbytes, &nrec_infos, &update_info, &hoss_rc);
if(ret == 0 && hoss_rc == 0 && nbytes == sizeof(ondisk_rep_factor))
{
@@ -339,7 +356,7 @@ static __blocking triton_ret_t resolve_replication_factor(
triton_debug(triton_dbg_rosd, "rosd_write() writing on-disk replication factor of %d\n", ondisk_rep_factor);
/* write new replication factor to disk */
- ret = hoss_write(&soid, 0, 1, sizeof(ondisk_rep_factor), 0, 0, 0, &ondisk_rep_factor, grp, &nbytes, &hoss_rc);
+ ret = hoss_write(&soid, 0, 1, sizeof(ondisk_rep_factor), HOSS_NONE, 0, 0, &ondisk_rep_factor, grp, &nbytes, &hoss_rc);
if(ret != 0 || hoss_rc != 0 || nbytes != sizeof(ondisk_rep_factor))
{
triton_error_msg("hoss_write() failure for replication factor, return code %d\n", ret);
@@ -379,6 +396,16 @@ static __blocking triton_ret_t triton_rpc_rosd_write(hg_handle_t handle)
triton_debug(triton_dbg_rosd, "rosd_write() with incoming replication factor of %d and incoming expected position of %d\n", in.replication_factor, in.expected_position);
+ /* argument checking */
+ if(in.container == 0 || in.object == 0)
+ {
+ out.tret = TRITON_ERR_INVAL;
+ triton_error_msg("Error: cannot create object in container or object 0.\n");
+ triton_mercury_respond(handle, &out);
+ HG_Destroy(handle);
+ return(TRITON_SUCCESS);
+ }
+
/* downstream replicas must have replication factor already defined for
* them.
*/
@@ -554,6 +581,7 @@ static __blocking triton_ret_t __remote_triton_rpc_rosd_write(
hg_handle_t handle = HG_HANDLE_NULL;
size_t bulk_size = recordcount*recordlen;
+ in.container = container;
in.object = object;
in.fork = fork;
in.start_record = start_record;
@@ -692,7 +720,7 @@ static __blocking triton_ret_t rosd_write_local_storage(
/* TODO: handle canceled error code (or whatever hoss will report to
* indicate that a txn failed and should be retried)
*/
- ret = hoss_begin(&grp, NULL, 0, g_hoss_ctx);
+ ret = hoss_begin(&grp, NULL, HOSS_MSYNC|HOSS_ORDERED|HOSS_DEFERRED, g_hoss_ctx);
if(ret != 0)
{
return(TRITON_ERR_IO);
@@ -705,8 +733,8 @@ static __blocking triton_ret_t rosd_write_local_storage(
soid_eids[2] = fork;
/* TODO: the update id arguments probably aren't right here... */
- ret = hoss_write(&soid, start_record, recordcount, recordlen, 0, 0, update_id_condition, buffer, grp, &nbytes, &hoss_rc);
- if(ret != 0 || hoss_rc != 0 || nbytes != recordcount*recordlen)
+ ret = hoss_write(&soid, start_record, recordcount, recordlen, HOSS_NONE, 0, update_id_condition, buffer, grp, &nbytes, &hoss_rc);
+ if(ret != 0)
{
triton_error_msg("hoss_write() failure, return code %d\n", ret);
hoss_end(grp, 0);
@@ -719,6 +747,14 @@ static __blocking triton_ret_t rosd_write_local_storage(
return(TRITON_ERR_IO);
}
+ /* NOTE: output arguments to hoss_write (rc and nbytes) are not filled
+ * in until completion of hoss_end()
+ */
+ if(hoss_rc != 0 || nbytes != recordcount*recordlen)
+ {
+ triton_error_msg("hoss_write() failure or short write.\n");
+ return(TRITON_ERR_IO);
+ }
return(TRITON_SUCCESS);
}
diff --git a/code/src/replicated-osd/rosd.ae b/code/src/replicated-osd/rosd.ae
index 4188ec5..d430b52 100644
--- a/code/src/replicated-osd/rosd.ae
+++ b/code/src/replicated-osd/rosd.ae
@@ -26,7 +26,7 @@ static int module_refcount = 0;
static int64_t rosd_read_max_allocation = -1;
static int64_t rosd_write_max_allocation = -1;
-static int64_t rosd_read_buffer_size = -1;
+int64_t rosd_read_buffer_size = -1;
int64_t rosd_write_buffer_size = -1;
uint32_t rosd_default_replication_factor = -1;
int64_t rosd_read_xfer_pipeline_depth = -1;
diff --git a/code/src/replicated-osd/rosd.hae b/code/src/replicated-osd/rosd.hae
index 9746d5f..58f2290 100644
--- a/code/src/replicated-osd/rosd.hae
+++ b/code/src/replicated-osd/rosd.hae
@@ -26,6 +26,11 @@
*/
#define ROSD_FLAG_FIRST_WRITE 16
+
+/* special values */
+/* TODO: how to sync these with ASG equivalents cleanly? */
+#define ROSD_UPDATE_ID_MIXED UINT64_MAX
+
/* TODO: we need to decide whether to just go with asg types or not - shadowing
* structs like asg_record_info_t with our own basically-equivalent types is
* silly and error-prone */
@@ -78,19 +83,29 @@ __blocking triton_ret_t remote_triton_rpc_rosd_write(
uint64_t *new_update_id,
uint32_t replication_factor);
-__blocking triton_ret_t remote_triton_rpc_rosd_read(
+__blocking triton_ret_t remote_triton_rpc_rosd_read_sequence(
uint64_t container,
uint64_t object,
uint64_t fork,
- uint64_t start_record,
- uint64_t recordcount,
+ uint64_t record_start,
+ uint64_t record_count,
+ uint64_t record_expected_len,
int flags,
+ uint64_t update_id_condition,
void * buf,
- uint64_t buf_size,
- uint64_t * version_info,
- uint64_t * transferred,
- struct rosd_record_info * record_info,
- uint64_t * record_info_xferred);
+ struct rosd_record_info * record_info);
+
+__blocking triton_ret_t remote_triton_rpc_rosd_read_one(
+ uint64_t container,
+ uint64_t object,
+ uint64_t fork,
+ uint64_t record,
+ uint32_t flags,
+ uint64_t update_id_condition,
+ void * buf,
+ size_t buf_size,
+ uint64_t record_len,
+ uint64_t * record_update_id);
__blocking triton_ret_t rosd_pipeline_setup_one_buffer(
struct buffer_mgmt_instance* buffer_instance,
@@ -133,12 +148,11 @@ __blocking triton_ret_t remote_triton_rpc_rosd_reset(
uint64_t container,
uint64_t object,
uint64_t fork,
- uint64_t record,
- uint64_t rcount,
- uint64_t rlen,
+ uint64_t record_start,
+ uint64_t record_count,
int flags,
uint64_t condition,
- uint64_t *reset_count);
+ uint64_t *records_reset);
__blocking triton_ret_t remote_triton_rpc_rosd_alloc(
na_addr_t addr,
diff --git a/code/src/server/triton-server.ae b/code/src/server/triton-server.ae
index 1653d91..fa12d4f 100644
--- a/code/src/server/triton-server.ae
+++ b/code/src/server/triton-server.ae
@@ -10,6 +10,8 @@
#include <aesop/aesop.h>
#include <hoss.hae>
+#include <recordstore.hae>
+#include <idb.h>
#include "src/remote/mercury-engine.hae"
#include "src/common/resources/signal/signal.hae"
#include "src/common/triton-debug.h"
@@ -107,11 +109,14 @@ __blocking int aesop_main(int argc, char **argv)
triton_ret_t tret;
char* listen_addr;
char* conffile;
- const char * hoss_path;
+ const char * idb_path;
+ const char * rs_opts;
char logname[PATH_MAX];
int daemonize;
+ IDBC idbc;
struct hoss_ctx *hoss_ctx;
int ret;
+ rs_instance_t rit;
parse_args(&argc,
&argv,
@@ -175,33 +180,51 @@ __blocking int aesop_main(int argc, char **argv)
triton_info_msg("Server starting\n");
/* set up local storage */
- /* TODO: REPLACEME
- tret = tosd_init();
- if(triton_is_error(tret))
+
+ rs_opts = triton_zeroconf_get("triton_rs_opts");
+ if(!rs_opts)
{
- triton_error_print(tret, "tosd_init()");
- triton_error_destroy(tret);
+ triton_error_msg("config variable triton_rs_opts "
+ "not present, exiting...\n");
+ triton_place_finalize();
+ triton_debug_finalize();
+ return(-1);
+ }
+ rit = rs_init("file", rs_opts);
+ if(!rit)
+ {
+ triton_error_msg_ret(tret, "rs_init()");
triton_place_finalize();
triton_debug_finalize();
return(-1);
}
- */
- /* TODO: better encapsulate this */
- hoss_path = triton_zeroconf_get("triton_hoss_base_path");
- if (hoss_path == NULL){
- triton_error_msg("config variable triton_hoss_base_path "
+ idb_path = triton_zeroconf_get("triton_idb_path");
+ if (idb_path == NULL){
+ triton_error_msg("config variable triton_idb_path "
"not present, exiting...\n");
- /* tosd_finalize(); */
+ rs_finalize(rit);
+ triton_place_finalize();
+ triton_debug_finalize();
+ return(-1);
+ }
+
+ idbc = idb_init(idb_path, IDB_SHARED_TABLE|IDB_USE_SNAPSHOT);
+ if(!idbc)
+ {
+ triton_error_msg_ret(tret, "rs_init()");
+ rs_finalize(rit);
triton_place_finalize();
triton_debug_finalize();
return(-1);
}
- ret = hoss_init(&hoss_ctx, hoss_path);
+
+ ret = hoss_init_config(&hoss_ctx, rit, idbc);
if(ret != 0)
{
triton_error_msg_ret(tret, "hoss_init()");
- /*tosd_finalize();*/
+ rs_finalize(rit);
+ idb_fini(idbc);
triton_place_finalize();
triton_debug_finalize();
return(-1);
@@ -220,7 +243,8 @@ __blocking int aesop_main(int argc, char **argv)
triton_error_msg_ret(tret, "triton_signal_init()");
triton_error_destroy(tret);
hoss_fini(hoss_ctx);
- tosd_finalize();
+ rs_finalize(rit);
+ idb_fini(idbc);
triton_place_finalize();
triton_debug_finalize();
return(-1);
@@ -236,7 +260,8 @@ __blocking int aesop_main(int argc, char **argv)
triton_error_msg_ret(tret, "triton_rosd_svr_init()");
triton_error_destroy(tret);
hoss_fini(hoss_ctx);
- /*tosd_finalize();*/
+ rs_finalize(rit);
+ idb_fini(idbc);
triton_place_finalize();
triton_debug_finalize();
#if 0
@@ -254,7 +279,8 @@ __blocking int aesop_main(int argc, char **argv)
triton_error_destroy(tret);
triton_rosd_svr_finalize();
hoss_fini(hoss_ctx);
- /*tosd_finalize();*/
+ rs_finalize(rit);
+ idb_fini(idbc);
#if 0
triton_signal_finalize();
#endif
@@ -272,7 +298,8 @@ __blocking int aesop_main(int argc, char **argv)
triton_mercury_engine_finalize();
triton_rosd_svr_finalize();
hoss_fini(hoss_ctx);
- /*tosd_finalize();*/
+ rs_finalize(rit);
+ idb_fini(idbc);
#if 0
triton_signal_finalize();
#endif
@@ -344,7 +371,8 @@ __blocking int aesop_main(int argc, char **argv)
triton_rosd_svr_finalize();
hoss_fini(hoss_ctx);
- /*tosd_finalize();*/
+ rs_finalize(rit);
+ idb_fini(idbc);
triton_core_rpc_finalize();
triton_mercury_engine_finalize();
#if 0
diff --git a/code/src/zeroconf/zeroconf.c b/code/src/zeroconf/zeroconf.c
index 8022b32..bb81d5a 100644
--- a/code/src/zeroconf/zeroconf.c
+++ b/code/src/zeroconf/zeroconf.c
@@ -40,10 +40,16 @@ struct zeroconf_entry zeroconf_array[] =
"Debug mask to enable debugging of various triton components",
NULL,
},
- /* hoss configuration parameters (TODO: grow) */
+ /* recordstore configuration parameters */
{
- CFG_STR("triton_hoss_base_path", 0, CFGF_NONE),
- "HOSS data/metadata path",
+ CFG_STR("triton_rs_opts", 0, CFGF_NONE),
+ "Configuration string to pass into recordstore, including path:<PATH>",
+ NULL
+ },
+ /* idb configuration parameters */
+ {
+ CFG_STR("triton_idb_path", 0, CFGF_NONE),
+ "IDB storage path",
NULL
},
/* object placement */
diff --git a/code/tests/Makefile.subdir b/code/tests/Makefile.subdir
index 570a3f9..183aedb 100644
--- a/code/tests/Makefile.subdir
+++ b/code/tests/Makefile.subdir
@@ -13,7 +13,8 @@ check_PROGRAMS += \
tests/placement/test-placement \
tests/system-state/test-system-state \
tests/asg/test-asg-simple \
- tests/asg/test-asg-multi-init
+ tests/asg/test-asg-multi-init \
+ tests/asg/test-asg-cond
#TODO: REPLACEME
#tests/transactional-osd/tosd1
@@ -40,6 +41,7 @@ TESTS += \
tests/common/test-debug.sh \
tests/system-state/test-system-state \
tests/asg/test-asg-simple.sh \
+ tests/asg/test-asg-cond.sh \
tests/asg/test-asg-multi-init.sh
#TODO: REPLACEME
@@ -58,6 +60,7 @@ AE_SRC += tests/placement/test-placement.ae \
tests_common_test_triton_ret_t_SOURCES = tests/common/test-triton-ret-t.c
tests_common_test_debug_SOURCES = tests/common/test-debug.c
tests_asg_test_asg_simple_SOURCES = tests/asg/test-asg-simple.c
+tests_asg_test_asg_cond_SOURCES = tests/asg/test-asg-cond.c
tests_asg_test_asg_multi_init_SOURCES = tests/asg/test-asg-multi-init.c
#tests_transactional_osd_bdb_snapshot_write_SOURCES = tests/transactional-osd/bdb-snapshot-write.c
diff --git a/code/tests/asg/test-asg-cond.c b/code/tests/asg/test-asg-cond.c
new file mode 100644
index 0000000..0e4c3a9
--- /dev/null
+++ b/code/tests/asg/test-asg-cond.c
@@ -0,0 +1,248 @@
+#include <stdio.h>
+#include <errno.h>
+#include <assert.h>
+#include <stdio.h>
+#include <unistd.h>
+#include <stdlib.h>
+
+#include "include/asg.h"
+
+// expects at least a non-empty string fmt
+#define ERR(...) \
+ do { \
+ fprintf(stderr, "%s:%d:", __FILE__, __LINE__); \
+ fprintf(stderr, __VA_ARGS__); \
+ return -1; \
+ } while (0)
+
+int main(int argc, char **argv)
+{
+ int ret;
+ asg_instance_t ait;
+ asg_update_id_t out_ver;
+ asg_container_id_t cid;
+ asg_object_id_t oid;
+ asg_fork_id_t fid;
+ asg_record_id_t rid;
+ asg_size_t out_size;
+ char fill = 'c';
+ asg_record_info_t rinfo;
+
+ if(argc != 2)
+ {
+ fprintf(stderr, "Usage: test-asg-simple <server address>\n");
+ fprintf(stderr, " example: ./test-asg-simple tcp://localhost:3344\n");
+ return(-1);
+ }
+
+ oid = 666;
+
+ cid = 1;
+ fid = 1;
+ rid = 1;
+
+ ret = asg_initialize(&ait, argv[1]);
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_initialize(..., \"%s\")\n", argv[1]);
+
+ out_ver = -1;
+ out_size = -1;
+
+ /* shorthand */
+#define LOC \
+ ait, ASG_LOCATION_AUTO, cid, oid, fid, rid
+
+ /* write data (effectively a create) */
+ ret = asg_write(
+ LOC,
+ 1, /* rec count */
+ 1, /* rec len */
+ ASG_COND_ALL, /* flags */
+ 1, /* update id */
+ &out_ver, /* assigned version */
+ &fill, /* data */
+ &out_size); /* resulting size */
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_write()\n");
+ else if (out_size != 1)
+ ERR("Error: asg_write() (short write)\n");
+ /* the update ID should be exactly what we gave it */
+ else if (out_ver != 1)
+ ERR("Error: gave ver 1 to asg_write, got %lu\n", out_ver);
+
+ out_ver = -1;
+ out_size = -1;
+ /* overwrite with a version ID = last - should fail */
+ ret = asg_write(
+ LOC,
+ 1,
+ 1,
+ ASG_COND_ALL,
+ 1,
+ &out_ver,
+ &fill,
+ &out_size);
+ if (ret != ASG_ERR_UPDATE_ID)
+ ERR("asg_write(): expected update id error\n");
+
+ memset(&rinfo, 0xff, sizeof(rinfo));
+ /* read with a version ID - should fail (ver is >=) */
+ ret = asg_read_sequence(
+ LOC,
+ 1, /* rec count */
+ 1, /* expected record length */
+ ASG_COND_ALL, /* flags */
+ 1, /* update id */
+ &fill,
+ &rinfo);
+ if(ret != ASG_ERR_UPDATE_ID)
+ ERR("asg_read_sequence(): expected update id error\n");
+ /* succeed (ver is >) */
+ ret = asg_read_sequence(
+ LOC,
+ 1, /* rec count */
+ 1, /* expected record length */
+ ASG_COND_ALL, /* flags */
+ 2, /* update id */
+ &fill,
+ &rinfo);
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_read_sequence()\n");
+ else if (rinfo.record_id != 1 ||
+ rinfo.seq_len != 1 || rinfo.record_len != 1)
+ ERR("Error: asg_read_sequence()\n");
+ else if (rinfo.record_update_id != 1)
+ ERR("Error: expected ver 1 from asg_read_sequence\n");
+
+ memset(&rinfo, 0xff, sizeof(rinfo));
+ /* read with a version ID - should fail (ver is >=) */
+ ret = asg_read_one(
+ LOC,
+ ASG_COND_ALL,
+ 1,
+ &fill,
+ 1,
+ &rinfo);
+ if (ret != ASG_ERR_UPDATE_ID)
+ ERR("asg_read_one(): expected update id error\n");
+ /* should succeed (ver is >) */
+ ret = asg_read_one(
+ LOC,
+ ASG_COND_ALL,
+ 1,
+ &fill,
+ 2,
+ &rinfo);
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_read_one()\n");
+ else if (rinfo.record_id != 1 ||
+ rinfo.seq_len != 1 || rinfo.record_len != 1)
+ ERR("Error: asg_read_one()\n");
+ else if (rinfo.record_update_id != 1)
+ ERR("Error: expected ver 1 from asg_read_one\n");
+
+ out_size = -1;
+ out_ver = -1;
+ /* now try to increment */
+ ret = asg_write(
+ LOC,
+ 1,
+ 1,
+ ASG_COND_ALL,
+ 2,
+ &out_ver,
+ &fill,
+ &out_size);
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_write()\n");
+ else if (out_size != 1)
+ ERR("Error: asg_write() (short write)\n");
+ /* the update ID should be exactly what we gave it */
+ else if (out_ver != 2)
+ ERR("Error: gave ver 2 to asg_write, got %lu\n", out_ver);
+
+ out_ver = -1; out_size = -1;
+ /* overwrite with a version ID <= last - should fail */
+ ret = asg_write(
+ LOC,
+ 1,
+ 1,
+ ASG_COND_ALL,
+ 2,
+ &out_ver,
+ &fill,
+ &out_size);
+ if (ret != ASG_ERR_UPDATE_ID)
+ ERR("asg_write(): expected update id error\n");
+
+ memset(&rinfo, 0xff, sizeof(rinfo));
+ /* read with a version ID - should fail (ver is >=) */
+ ret = asg_read_sequence(
+ LOC,
+ 1, /* rec count */
+ 1, /* expected record length */
+ ASG_COND_ALL, /* flags */
+ 2, /* update id */
+ &fill,
+ &rinfo);
+ if(ret != ASG_ERR_UPDATE_ID)
+ ERR("asg_read_sequence(): expected update id error\n");
+ /* succeed (ver is >) */
+ ret = asg_read_sequence(
+ LOC,
+ 1, /* rec count */
+ 1, /* expected record length */
+ ASG_COND_ALL, /* flags */
+ 3, /* update id */
+ &fill,
+ &rinfo);
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_read_sequence()\n");
+ else if (rinfo.record_id != 1 ||
+ rinfo.seq_len != 1 || rinfo.record_len != 1)
+ ERR("Error: asg_read_one()\n");
+ else if (rinfo.record_update_id != 2)
+ ERR("Error: expected ver 2 from asg_read_sequence\n");
+
+ memset(&rinfo, 0xff, sizeof(rinfo));
+ /* read with a version ID - should fail (ver is >=) */
+ ret = asg_read_one(
+ LOC,
+ ASG_COND_ALL,
+ 1,
+ &fill,
+ 2,
+ &rinfo);
+ if (ret != ASG_ERR_UPDATE_ID)
+ ERR("asg_read_one(): expected update id error\n");
+ /* should succeed (ver is >) */
+ ret = asg_read_one(
+ LOC,
+ ASG_COND_ALL,
+ 1,
+ &fill,
+ 3,
+ &rinfo);
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_read_one()\n");
+ else if (rinfo.record_id != 1 ||
+ rinfo.seq_len != 1 || rinfo.record_len != 1)
+ ERR("Error: asg_read_one()\n");
+ else if (rinfo.record_update_id != 1)
+ ERR("Error: expected ver 1 from asg_read_one\n");
+
+ ret = asg_finalize(ait);
+ if(ret != ASG_SUCCESS)
+ ERR("Error: asg_finalize()\n");
+
+ return(0);
+}
+
+/*
+ * Local variables:
+ * c-indent-level: 4
+ * c-basic-offset: 4
+ * End:
+ *
+ * vim: ft=c ts=8 sts=4 sw=4 expandtab
+ */
diff --git a/code/tests/asg/test-asg-simple.sh b/code/tests/asg/test-asg-cond.sh
similarity index 91%
copy from code/tests/asg/test-asg-simple.sh
copy to code/tests/asg/test-asg-cond.sh
index 18af42f..57cb9b8 100755
--- a/code/tests/asg/test-asg-simple.sh
+++ b/code/tests/asg/test-asg-cond.sh
@@ -11,7 +11,7 @@ test_start_servers 1 5 60 1
# actual test case
#####################
-run_to 60 tests/asg/test-asg-simple $svr1 555
+run_to 60 tests/asg/test-asg-cond $svr1
if [ $? -ne 0 ]; then
run_to 60 src/admin-tools/triton-shutdown-all-servers $svr1 &> /dev/null
wait
diff --git a/code/tests/asg/test-asg-simple.c b/code/tests/asg/test-asg-simple.c
index 9dbf4c0..4001e81 100644
--- a/code/tests/asg/test-asg-simple.c
+++ b/code/tests/asg/test-asg-simple.c
@@ -24,6 +24,7 @@ int main(int argc, char **argv)
char fill = '\0';
asg_fork_info_t fork_buf;
asg_fork_id_t next_fid;
+ asg_record_info_t rinfo;
if(argc != 3)
{
@@ -93,7 +94,7 @@ int main(int argc, char **argv)
memset(buffer, 0, buffer_sz);
/* read data back out */
- ret = asg_read(
+ ret = asg_read_sequence(
ait,
ASG_LOCATION_AUTO,
cid, /* container */
@@ -101,17 +102,15 @@ int main(int argc, char **argv)
fid, /* fork */
rid, /* start rec */
buffer_sz, /* rec count */
+ 1, /* expected record length */
ASG_COND_NONE, /* flags */
0, /* update id */
buffer,
- buffer_sz,
- &out_size, /* resulting size */
- NULL, /* record info */
- NULL, /* number of record info bufs */
- &out_ver); /* assigned version */
+ &rinfo);
+
if(ret != ASG_SUCCESS)
{
- fprintf(stderr, "Error: asg_read()\n");
+ fprintf(stderr, "Error: asg_read_sequence()\n");
return(-1);
}
@@ -126,9 +125,137 @@ int main(int argc, char **argv)
}
fill++;
}
+ memset(buffer, 0, buffer_sz);
+
+ /* now simply read an arbitrary byte */
+ ret = asg_read_one(
+ ait,
+ ASG_LOCATION_AUTO,
+ cid,
+ oid,
+ fid,
+ buffer_sz / 2,
+ ASG_COND_NONE,
+ 0,
+ buffer,
+ buffer_sz,
+ &rinfo);
+ if (ret != ASG_SUCCESS) {
+ fprintf(stderr, "Error: asg_read_one()\n");
+ return -1;
+ }
+
+ fill = '\0';
+ for (i = 0; i < buffer_sz/2; i++)
+ fill++;
+ if (buffer[i] != fill) {
+ fprintf(stderr,
+ "Error: buffer validation failure (asg_read_one()) "
+ "starting at offset %d\n", i);
+ return -1;
+ }
+
+ /* try to read a nonexistent record */
+ ret = asg_read_one(
+ ait,
+ ASG_LOCATION_AUTO,
+ cid,
+ oid,
+ fid+1,
+ 1,
+ ASG_COND_NONE,
+ 0,
+ buffer,
+ buffer_sz,
+ &rinfo);
+ if (ret != ASG_ERR_NOENT) {
+ fprintf(stderr,
+ "Error: expected noent return (%d), got %d (asg_read_one)\n",
+ ASG_ERR_NOENT, ret);
+ return -1;
+ }
+
+ fill = '\0';
+ for (i = 0; i < buffer_sz; i++)
+ buffer[i] = fill++;
+
+ /* write a multi-byte record */
+ ret = asg_write(
+ ait,
+ ASG_LOCATION_AUTO,
+ cid,
+ oid,
+ fid+1,
+ 1, /* record */
+ 1, /* rec count */
+ buffer_sz, /* rec len */
+ ASG_COND_NONE,
+ 0,
+ &out_ver,
+ buffer,
+ &out_size);
+ if (ret != ASG_SUCCESS || out_size != buffer_sz) {
+ fprintf(stderr, "Error: asg_write()\n");
+ return -1;
+ }
+
+ /* read with insufficient buffer space */
+ ret = asg_read_one(
+ ait,
+ ASG_LOCATION_AUTO,
+ cid,
+ oid,
+ fid+1,
+ 1,
+ ASG_COND_NONE,
+ 0,
+ buffer,
+ buffer_sz/2,
+ &rinfo);
+ if (ret != ASG_ERR_BUFFER_SMALL) {
+ fprintf(stderr,
+ "Error: expected bufsmall return (%d), "
+ "got %d (asg_read_one)\n",
+ ASG_ERR_BUFFER_SMALL, ret);
+ return -1;
+ }
+
+ /* read with sufficient buffer space */
+ ret = asg_read_one(
+ ait,
+ ASG_LOCATION_AUTO,
+ cid,
+ oid,
+ fid+1,
+ 1,
+ ASG_COND_NONE,
+ 0,
+ buffer,
+ buffer_sz,
+ &rinfo);
+ if (ret != ASG_SUCCESS) {
+ fprintf(stderr, "Error: asg_read_one()\n");
+ return -1;
+ }
+
+ /* check buffer */
+ fill = '\0';
+ for(i=0; i<buffer_sz; i++)
+ {
+ if (buffer[i] != fill) {
+ fprintf(stderr,
+ "Error: buffer validation failure (asg_read_one()) "
+ "starting at offset %d\n", i);
+ return -1;
+ }
+ fill++;
+ }
+
free(buffer);
+ /* TODO: re-enable remainder of test once functions are working */
+#if 0
ret = asg_probe_object(ait,
ASG_LOCATION_AUTO,
cid,
@@ -169,7 +296,8 @@ int main(int argc, char **argv)
&count,
&next_fid);
assert(count == 0);
-
+#endif
+
ret = asg_finalize(ait);
if(ret != ASG_SUCCESS)
{
diff --git a/code/tests/placement/test-placement.ae b/code/tests/placement/test-placement.ae
index caf9d95..3b97a8a 100644
--- a/code/tests/placement/test-placement.ae
+++ b/code/tests/placement/test-placement.ae
@@ -70,7 +70,7 @@ __blocking int aesop_main(int argc, char **argv)
oid_l.l = oid;
oid_l.u = 0;
- tret = triton_place_lookup(oid, closest, 3);
+ tret = triton_place_lookup(oid, closest, 3, NULL);
if(triton_is_error(tret))
{
triton_error_print(tret, "triton_place_lookup()");
@@ -102,7 +102,7 @@ __blocking int aesop_main(int argc, char **argv)
oid -= 3;
oid_l.l = oid;
- tret = triton_place_lookup(oid, closest, 3);
+ tret = triton_place_lookup(oid, closest, 3, NULL);
if(triton_is_error(tret))
{
triton_error_print(tret, "triton_place_lookup()");
@@ -143,7 +143,7 @@ __blocking int aesop_main(int argc, char **argv)
oid = servers[3];
oid.u += 3;
- tret = triton_place_lookup(oid, closest, 3);
+ tret = triton_place_lookup(oid, closest, 3, NULL);
if(triton_is_error(tret))
{
triton_error_print(tret, "triton_place_lookup()");
diff --git a/code/tests/test-util.sh b/code/tests/test-util.sh
index 40ec7e9..ceca307 100644
--- a/code/tests/test-util.sh
+++ b/code/tests/test-util.sh
@@ -40,7 +40,8 @@ function test_gen_conf ()
done
echo "}" >> ${TMPBASE}/triton-$pid.conf
echo "triton_debug_file = ${TMPBASE}/triton-server-%TRITON_SERVER_NAME%.log" >> ${TMPBASE}/triton-$pid.conf
- echo "triton_sos_base_path = ${TMPBASE}/triton-server-%TRITON_SERVER_NAME%-sos" >> ${TMPBASE}/triton-$pid.conf
+ echo "triton_idb_path = ${TMPBASE}/triton-server-%TRITON_SERVER_NAME%-storage" >> ${TMPBASE}/triton-$pid.conf
+ echo "triton_rs_opts = \"key_prefix_bytes:24,path:${TMPBASE}/triton-server-%TRITON_SERVER_NAME%-storage\"" >> ${TMPBASE}/triton-$pid.conf
echo "triton_debug_masks = \"all\"" >> ${TMPBASE}/triton-$pid.conf
if [[ $repfactor -gt 0 ]] ; then
echo "triton_rosd_default_replication_factor = $repfactor" \
hooks/post-receive
--
1
0
24 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 2a74c490ebf926662e046834c1f30d234b9e37e6 (commit)
from af2d0f03aaf0f3af7ab2f1335cf86868f8b700cb (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 2a74c490ebf926662e046834c1f30d234b9e37e6
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Tue Feb 24 16:33:05 2015 -0500
add hoss perf graph
-----------------------------------------------------------------------
Summary of changes:
.../figs/hoss-ssd/dd-1way.txt | 13 ++++++++++
.../figs/hoss-ssd/hoss-16way.txt | 13 ++++++++++
.../figs/hoss-ssd/hoss-1way.txt | 13 ++++++++++
.../figs/hoss-ssd/hoss-ssd.pdf | Bin 0 -> 8200 bytes
.../recordstore-ssd.plt => hoss-ssd/hoss-ssd.plt} | 6 ++--
.../localstore-report.tex | 24 ++++++++++++-------
6 files changed, 57 insertions(+), 12 deletions(-)
create mode 100644 reports/triton-localstore-convergence/figs/hoss-ssd/dd-1way.txt
create mode 100644 reports/triton-localstore-convergence/figs/hoss-ssd/hoss-16way.txt
create mode 100644 reports/triton-localstore-convergence/figs/hoss-ssd/hoss-1way.txt
create mode 100644 reports/triton-localstore-convergence/figs/hoss-ssd/hoss-ssd.pdf
copy reports/triton-localstore-convergence/figs/{recordstore-ssd/recordstore-ssd.plt => hoss-ssd/hoss-ssd.plt} (77%)
Diff of changes:
diff --git a/reports/triton-localstore-convergence/figs/hoss-ssd/dd-1way.txt b/reports/triton-localstore-convergence/figs/hoss-ssd/dd-1way.txt
new file mode 100644
index 0000000..27375d0
--- /dev/null
+++ b/reports/triton-localstore-convergence/figs/hoss-ssd/dd-1way.txt
@@ -0,0 +1,13 @@
+# dd if=/dev/zero of=/tmp/1.dat bs=SIZE count=50000 oflag=direct,dsync
+# <size> <MiB/s>
+4096 7.73632812500000000000
+8192 15.10546875000000000000
+16384 29.66406250000000000000
+32768 54.07812500000000000000
+65536 93.32500000000000000000
+131072 148.93750000000000000000
+262144 178.35000000000000000000
+524288 253.85000000000000000000
+1048576 261.40000000000000000000
+2097152 268.00000000000000000000
+4194304 256.00000000000000000000
diff --git a/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-16way.txt b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-16way.txt
new file mode 100644
index 0000000..710e304
--- /dev/null
+++ b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-16way.txt
@@ -0,0 +1,13 @@
+# ./hoss-concurrent-io SIZE 16 10 alignment:4096,O_SYNC,O_DIRECT,sync_highwater:8,path:/tmp/recordstore
+# <size> <MiB/s>
+4096 19.234492
+8192 39.090148
+16384 90.651492
+32768 140.364090
+65536 222.935439
+131072 257.488435
+262144 245.482159
+524288 239.529513
+1048576 288.555086
+2097152 289.807723
+4194304 290.419818
diff --git a/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-1way.txt b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-1way.txt
new file mode 100644
index 0000000..940b420
--- /dev/null
+++ b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-1way.txt
@@ -0,0 +1,13 @@
+# ./hoss-concurrent-io SIZE 1 10 alignment:4096,O_SYNC,O_DIRECT,sync_highwater:8,path:/tmp/recordstore
+# <size> <MiB/s>
+4096 4.822872
+8192 9.473515
+16384 18.572041
+32768 37.385085
+65536 62.590583
+131072 94.291860
+262144 132.446997
+524288 186.848998
+1048576 207.131111
+2097152 205.830716
+4194304 188.557519
diff --git a/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-ssd.pdf b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-ssd.pdf
new file mode 100644
index 0000000..e9b6820
Binary files /dev/null and b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-ssd.pdf differ
diff --git a/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.plt b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-ssd.plt
similarity index 77%
copy from reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.plt
copy to reports/triton-localstore-convergence/figs/hoss-ssd/hoss-ssd.plt
index 6c3d2d7..459edde 100644
--- a/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.plt
+++ b/reports/triton-localstore-convergence/figs/hoss-ssd/hoss-ssd.plt
@@ -9,7 +9,7 @@ set style line 4 lc 5 lw 2
set style line 5 lc 2 lw 2
set style increment user
-set output "recordstore-ssd.eps"
+set output "hoss-ssd.eps"
set xlabel "Access size (bytes)"
set ylabel "Bandwidth (MiB/s)" offset 1,0
@@ -24,7 +24,7 @@ set rmargin 5
set xtics ("4 KiB" 4096, "8 KiB" 8192, "16 KiB" 16384, "32 KiB" 32768, "64 KiB" 65536, "128 KiB" 131072, "256 KiB" 262144, "512 KiB" 524288, "1 MiB" 1048576, "2 MiB" 2097152, "4 MiB" 4194304)
# set yrange [0:2100]
-plot "rs-16way.txt" using 1:2 with linespoints lt 1 pt 1 linecolor rgb "blue" title "RS (16x concurrency)", \
-"rs-1way.txt" using 1:2 with linespoints lt 1 pt 2 linecolor rgb "green" title "RS (no concurrency)", \
+plot "hoss-16way.txt" using 1:2 with linespoints lt 1 pt 1 linecolor rgb "blue" title "HOSS (16x concurrency)", \
+"hoss-1way.txt" using 1:2 with linespoints lt 1 pt 2 linecolor rgb "green" title "HOSS (no concurrency)", \
"dd-1way.txt" using 1:2 with linespoints lt 1 pt 3 linecolor rgb "orange" title "dd"
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index 422fc45..ead5f41 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -3,6 +3,7 @@
\usepackage{listings}
\usepackage{appendix}
\usepackage{color}
+\usepackage[tight,footnotesize]{subfigure}
\usepackage[pdftex]{graphicx}
\begin{document}
\title{Convergence on a shared API for local storage abstraction\\Deliverable Report 2.3.6a}
@@ -180,19 +181,24 @@ stored in Triton.
\subsection{Evaluation}
\begin{figure}[t]
-\centering
-\includegraphics[width=0.5\textwidth]{figs/recordstore-ssd/recordstore-ssd.pdf}
-\caption{Comparison of Recordstore (RS) write bandwidth and \texttt{dd}
-bandwidth on an Intel 730 series SSD using directio and synchronize I/O
-operations.}
-\label{fig:rs-perf}
+ \centering
+ \subfigure[Recordstore (RS)]{
+ \centering
+ \includegraphics[width=0.45\textwidth]{figs/recordstore-ssd/recordstore-ssd.pdf}
+ \label{fig:rs-perf}
+ }
+ \subfigure[HOSS]{
+ \centering
+ \includegraphics[width=0.45\textwidth]{figs/hoss-ssd/hoss-ssd.pdf}
+ \label{fig:hoss-perf}
+ }
+ \caption{Streaming write performance at the Recordstore and HOSS component level on an Intel 730 series SSD using direct I/O and
+synchronizing each operation to disk.}
\end{figure}
-Figure~\ref{fig:rs-perf} shows \textcolor{red}{TODO: fill this in.
+Figure~\ref{fig:rs-perf} shows something and Figure~\ref{fig:hoss-perf} shows something else. \textcolor{red}{TODO: fill this in.
Preliminary results. See .txt files for command line details.}
-\textcolor{red}{TODO: same experiment at HOSS level.}
-
\section{Conclusion}
The current Triton prototype operating atop the HOSS local storage component
hooks/post-receive
--
1
0
24 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via af2d0f03aaf0f3af7ab2f1335cf86868f8b700cb (commit)
from 05d9b2c1dce6034a3c716fe38b126d5d06e65a97 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit af2d0f03aaf0f3af7ab2f1335cf86868f8b700cb
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Tue Feb 24 09:41:18 2015 -0500
preliminary perf results at recordstore level
-----------------------------------------------------------------------
Summary of changes:
.../figs/recordstore-ssd/dd-1way.txt | 13 ++++++++
.../figs/recordstore-ssd/recordstore-ssd.pdf | Bin 0 -> 8185 bytes
.../figs/recordstore-ssd/recordstore-ssd.plt | 30 ++++++++++++++++++++
.../figs/recordstore-ssd/rs-16way.txt | 13 ++++++++
.../figs/recordstore-ssd/rs-1way.txt | 13 ++++++++
.../localstore-report.tex | 16 ++++++++++
6 files changed, 85 insertions(+), 0 deletions(-)
create mode 100644 reports/triton-localstore-convergence/figs/recordstore-ssd/dd-1way.txt
create mode 100644 reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.pdf
create mode 100644 reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.plt
create mode 100644 reports/triton-localstore-convergence/figs/recordstore-ssd/rs-16way.txt
create mode 100644 reports/triton-localstore-convergence/figs/recordstore-ssd/rs-1way.txt
Diff of changes:
diff --git a/reports/triton-localstore-convergence/figs/recordstore-ssd/dd-1way.txt b/reports/triton-localstore-convergence/figs/recordstore-ssd/dd-1way.txt
new file mode 100644
index 0000000..86cdf8a
--- /dev/null
+++ b/reports/triton-localstore-convergence/figs/recordstore-ssd/dd-1way.txt
@@ -0,0 +1,13 @@
+# dd if=/dev/zero of=/tmp/1.dat bs=SIZE count=50000 oflag=direct,dsync
+# <size> <MiB/s>
+4096 7.52109375000000000000
+8192 14.95000000000000000000
+16384 29.09687500000000000000
+32768 54.39687500000000000000
+65536 95.60625000000000000000
+131072 151.43750000000000000000
+262144 214.60000000000000000000
+524288 208.95000000000000000000
+1048576 234.70000000000000000000
+2097152 256.00000000000000000000
+4194304 240.40000000000000000000
diff --git a/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.pdf b/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.pdf
new file mode 100644
index 0000000..7a91741
Binary files /dev/null and b/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.pdf differ
diff --git a/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.plt b/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.plt
new file mode 100644
index 0000000..6c3d2d7
--- /dev/null
+++ b/reports/triton-localstore-convergence/figs/recordstore-ssd/recordstore-ssd.plt
@@ -0,0 +1,30 @@
+#set term po eps color solid "NimbusSanL-Regu" 16 fontfile "/usr/share/texmf-texlive/fonts/type1/urw/helvetic/uhvr8a.pfb"
+set term po eps color dashed 18 size 4in,3in
+
+# color blindness work around
+set style line 1 lc 1 lw 2 ps 1
+set style line 2 lc 3 lw 2 ps 1
+set style line 3 lc 4 lw 2 ps 1
+set style line 4 lc 5 lw 2
+set style line 5 lc 2 lw 2
+set style increment user
+
+set output "recordstore-ssd.eps"
+
+set xlabel "Access size (bytes)"
+set ylabel "Bandwidth (MiB/s)" offset 1,0
+set key bottom right
+#set grid
+#set rmargin 2
+
+set logscale x
+set xtic nomirror rotate by -45
+set rmargin 5
+#set bmargin 8
+set xtics ("4 KiB" 4096, "8 KiB" 8192, "16 KiB" 16384, "32 KiB" 32768, "64 KiB" 65536, "128 KiB" 131072, "256 KiB" 262144, "512 KiB" 524288, "1 MiB" 1048576, "2 MiB" 2097152, "4 MiB" 4194304)
+
+# set yrange [0:2100]
+plot "rs-16way.txt" using 1:2 with linespoints lt 1 pt 1 linecolor rgb "blue" title "RS (16x concurrency)", \
+"rs-1way.txt" using 1:2 with linespoints lt 1 pt 2 linecolor rgb "green" title "RS (no concurrency)", \
+"dd-1way.txt" using 1:2 with linespoints lt 1 pt 3 linecolor rgb "orange" title "dd"
+
diff --git a/reports/triton-localstore-convergence/figs/recordstore-ssd/rs-16way.txt b/reports/triton-localstore-convergence/figs/recordstore-ssd/rs-16way.txt
new file mode 100644
index 0000000..6e499e5
--- /dev/null
+++ b/reports/triton-localstore-convergence/figs/recordstore-ssd/rs-16way.txt
@@ -0,0 +1,13 @@
+# ./recordstore-bench-write file alignment:4096,O_SYNC,O_DIRECT,sync_highwater:8,path:/tmp/recordstore SIZE 16 10 foo.txt
+# <size> <MiB/s>
+4096 27.466946
+8192 41.815755
+16384 88.811613
+32768 140.231684
+65536 204.578464
+131072 235.267898
+262144 269.450455
+524288 274.282007
+1048576 289.935528
+2097152 271.690400
+4194304 290.580264
diff --git a/reports/triton-localstore-convergence/figs/recordstore-ssd/rs-1way.txt b/reports/triton-localstore-convergence/figs/recordstore-ssd/rs-1way.txt
new file mode 100644
index 0000000..f1701dc
--- /dev/null
+++ b/reports/triton-localstore-convergence/figs/recordstore-ssd/rs-1way.txt
@@ -0,0 +1,13 @@
+# ./recordstore-bench-write file alignment:4096,O_SYNC,O_DIRECT,sync_highwater:8,path:/tmp/recordstore SIZE 1 10 foo.txt
+# <size> <MiB/s>
+4096 6.663482
+8192 13.023758
+16384 26.189085
+32768 47.217648
+65536 84.426165
+131072 118.032200
+262144 162.351595
+524288 193.244971
+1048576 200.811967
+2097152 212.938592
+4194304 194.483911
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index 92a34fb..422fc45 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -177,6 +177,22 @@ with Sandia to continue to improve the HOSS implementation and exand it to
include more functionality for querying and managing data once it has been
stored in Triton.
+\subsection{Evaluation}
+
+\begin{figure}[t]
+\centering
+\includegraphics[width=0.5\textwidth]{figs/recordstore-ssd/recordstore-ssd.pdf}
+\caption{Comparison of Recordstore (RS) write bandwidth and \texttt{dd}
+bandwidth on an Intel 730 series SSD using directio and synchronize I/O
+operations.}
+\label{fig:rs-perf}
+\end{figure}
+
+Figure~\ref{fig:rs-perf} shows \textcolor{red}{TODO: fill this in.
+Preliminary results. See .txt files for command line details.}
+
+\textcolor{red}{TODO: same experiment at HOSS level.}
+
\section{Conclusion}
The current Triton prototype operating atop the HOSS local storage component
hooks/post-receive
--
1
0
See <https://jenkins-ci.mcs.anl.gov/job/triton/180/>
------------------------------------------
Started by an SCM change
[EnvInject] - Loading node environment variables.
Building remotely on trounce.mcs.anl.gov in workspace <https://jenkins-ci.mcs.anl.gov/job/triton/ws/>
Deleting project workspace...
done
Checkout:triton / <https://jenkins-ci.mcs.anl.gov/job/triton/ws/> - hudson.remoting.Channel@6c1344da:trounce.mcs.anl.gov
Using strategy: Default
Last Built Revision: Revision 9c9d7b4da56fbac8bd3b1258198026ec6d0d91e2 (origin/master)
Cloning the remote Git repository
Cloning repository [email protected]:triton
git --version
git version 1.7.9.5
Fetching upstream changes from origin
Commencing build of Revision 9c9d7b4da56fbac8bd3b1258198026ec6d0d91e2 (origin/master)
Checking out Revision 9c9d7b4da56fbac8bd3b1258198026ec6d0d91e2 (origin/master)
[triton] $ /bin/sh -xe /tmp/hudson674172220491034471.sh
+ code/scripts/jenkins/build-jenkins.sh
code
misc
sim
=======================================================================
= Build 180 2015-02-22_00-03-03
= on trounce.mcs.anl.gov [mcs trounce.mcs.anl.gov]
=======================================================================
Triton source in <https://jenkins-ci.mcs.anl.gov/job/triton/ws/>
Building in <https://jenkins-ci.mcs.anl.gov/job/triton/ws/build>
====================================================================
===== INFO =========================================================
====================================================================
Triton Source tree : <https://jenkins-ci.mcs.anl.gov/job/triton/ws/>
Workspace : <https://jenkins-ci.mcs.anl.gov/job/triton/ws/build>
Cachedir : /tmp/triton-cache
=======================================================================
==== STEP 1: Fetch build dependencies =================================
=======================================================================
openpa.tar.gz: (found in cache) OK (3ad998bb26ac84ee7de262db94dd7656)
bmi.tar.gz: (found in cache) OK (82a8efe604a80c80d18dab76dfc13c38)
boost.tar.bz2: (found in cache) OK (d6eef4b4cacb2183f2bf265a5a03a354)
c-utils.tar.gz: (cached but invalid. Removing) (fetching) Build step 'Execute shell' marked build as failure
TAP Reports Processing: START
Looking for TAP results report in workspace using pattern: build/tritonbuild/tests/results.tap
Did not find any matching files.
1
1
20 Feb '15
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 05d9b2c1dce6034a3c716fe38b126d5d06e65a97 (commit)
from 2a1928b83c6db071d468958ea1bda0790373bbd3 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 05d9b2c1dce6034a3c716fe38b126d5d06e65a97
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Fri Feb 20 14:37:08 2015 -0500
brief conclusion
-----------------------------------------------------------------------
Summary of changes:
.../localstore-report.tex | 6 +++++-
1 files changed, 5 insertions(+), 1 deletions(-)
Diff of changes:
diff --git a/reports/triton-localstore-convergence/localstore-report.tex b/reports/triton-localstore-convergence/localstore-report.tex
index 3af6ffe..92a34fb 100644
--- a/reports/triton-localstore-convergence/localstore-report.tex
+++ b/reports/triton-localstore-convergence/localstore-report.tex
@@ -179,6 +179,10 @@ stored in Triton.
\section{Conclusion}
-TODO: why this stuff is relevant
+The current Triton prototype operating atop the HOSS local storage component
+has demonstrated that we are able to share the same local storage
+implementation as is used by the Sirocco team at Sandia. This will allow us
+to pool research, development, testing, and maintenance effort for this
+component of our respective storage systems moving forward.
\end{document}
hooks/post-receive
--
1
0