Triton-commits
Threads by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
August 2014
- 1 participants
- 68 discussions
30 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 3d26932b4156a963c1b2f0540d8244c2cab824b5 (commit)
via aed2e68e36886fc5015a5ceddc8ed738fd3bde5e (commit)
via efec8b05d2f17eca459a5af23b628172f24eb6e9 (commit)
via 593d17409c15609c101cd73b540df78b0a58d7fc (commit)
from 0363f97f87cf40e75f4ab55c4dd96c5ced648300 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 3d26932b4156a963c1b2f0540d8244c2cab824b5
Merge: aed2e68e36886fc5015a5ceddc8ed738fd3bde5e 0363f97f87cf40e75f4ab55c4dd96c5ced648300
Author: Dai Dong <dong.dai(a)ttu.edu>
Date: Sat Aug 30 10:40:27 2014 -0500
final version
commit aed2e68e36886fc5015a5ceddc8ed738fd3bde5e
Author: Dai Dong <dong.dai(a)ttu.edu>
Date: Sat Aug 30 10:36:32 2014 -0500
final version
commit efec8b05d2f17eca459a5af23b628172f24eb6e9
Merge: 593d17409c15609c101cd73b540df78b0a58d7fc 2403dd48604f971395168ca05dd962d64b1e3814
Author: Dai Dong <dong.dai(a)ttu.edu>
Date: Fri Aug 22 10:09:25 2014 -0500
Merge branch 'master' of git.mcs.anl.gov:triton-private
commit 593d17409c15609c101cd73b540df78b0a58d7fc
Author: Dai Dong <dong.dai(a)ttu.edu>
Date: Fri Aug 22 10:06:10 2014 -0500
revision for clearance
-----------------------------------------------------------------------
Summary of changes:
papers/meta-graph/abstract.tex | 2 +-
papers/meta-graph/ack.tex | 4 +
papers/meta-graph/bib.bib | 54 +++-
papers/meta-graph/conclusion.tex | 2 +-
papers/meta-graph/design.tex | 10 +-
.../meta-graph/exps/2013-graph.numbers/Index.zip | Bin 49177 -> 49188 bytes
.../2013-graph.numbers/Metadata/Properties.plist | Bin 340 -> 340 bytes
papers/meta-graph/exps/fdegree.pdf | Bin 102154 -> 102149 bytes
papers/meta-graph/exps/jdegree.pdf | Bin 31231 -> 31226 bytes
papers/meta-graph/exps/pdegree.pdf | Bin 120995 -> 121000 bytes
papers/meta-graph/exps/udegree.pdf | Bin 5067 -> 5073 bytes
papers/meta-graph/intro.tex | 16 +-
papers/meta-graph/main.bbl | 222 +++++++++++++
papers/meta-graph/main.blg | 56 ++++
papers/meta-graph/main.tex | 13 +-
papers/meta-graph/model.tex | 26 +-
papers/meta-graph/proto.tex | 50 ++--
papers/meta-graph/{ => report}/exps/fdegree.pdf | Bin 102154 -> 102149 bytes
.../{exps/udegree.pdf => report/exps/histplot.pdf} | Bin 5067 -> 6802 bytes
papers/meta-graph/{ => report}/exps/jdegree.pdf | Bin 31231 -> 31226 bytes
papers/meta-graph/{ => report}/exps/pdegree.pdf | Bin 120995 -> 121000 bytes
papers/meta-graph/report/exps/poweRlaw-cdf.pdf | Bin 0 -> 80052 bytes
papers/meta-graph/{ => report}/exps/udegree.pdf | Bin 5067 -> 5073 bytes
.../udegree.pdf => report/exps/user-histogram.pdf} | Bin 5067 -> 6178 bytes
papers/meta-graph/report/power-law-dist.aux | 32 ++
papers/meta-graph/report/power-law-dist.log | 337 ++++++++++++++++++++
papers/meta-graph/report/power-law-dist.pdf | Bin 0 -> 433924 bytes
papers/meta-graph/report/power-law-dist.synctex.gz | Bin 0 -> 43374 bytes
papers/meta-graph/report/power-law-dist.tex | 155 +++++++++
.../report}/usenix.sty | 0
papers/meta-graph/rscripts/fdegree.r | 8 +-
papers/meta-graph/rscripts/jdegree.r | 7 +-
papers/meta-graph/rscripts/pdegree.r | 8 +-
papers/meta-graph/rscripts/udegree.r | 7 +-
papers/meta-graph/usecases.tex | 18 +-
35 files changed, 946 insertions(+), 81 deletions(-)
create mode 100644 papers/meta-graph/main.bbl
create mode 100644 papers/meta-graph/main.blg
copy papers/meta-graph/{ => report}/exps/fdegree.pdf (95%)
copy papers/meta-graph/{exps/udegree.pdf => report/exps/histplot.pdf} (52%)
copy papers/meta-graph/{ => report}/exps/jdegree.pdf (80%)
copy papers/meta-graph/{ => report}/exps/pdegree.pdf (97%)
create mode 100644 papers/meta-graph/report/exps/poweRlaw-cdf.pdf
copy papers/meta-graph/{ => report}/exps/udegree.pdf (75%)
copy papers/meta-graph/{exps/udegree.pdf => report/exps/user-histogram.pdf} (57%)
create mode 100644 papers/meta-graph/report/power-law-dist.aux
create mode 100644 papers/meta-graph/report/power-law-dist.log
create mode 100644 papers/meta-graph/report/power-law-dist.pdf
create mode 100644 papers/meta-graph/report/power-law-dist.synctex.gz
create mode 100644 papers/meta-graph/report/power-law-dist.tex
copy papers/{triton-wire-protocol => meta-graph/report}/usenix.sty (100%)
Diff of changes:
diff --git a/papers/meta-graph/abstract.tex b/papers/meta-graph/abstract.tex
index 9dae760..4891914 100644
--- a/papers/meta-graph/abstract.tex
+++ b/papers/meta-graph/abstract.tex
@@ -1,6 +1,6 @@
\begin{abstract}
-HPC platforms are capable of generating huge amounts of metadata about different entities including jobs, users, and files etc. \textit{Simple metadata}, which describe the attributes of these entities has already been well recorded and used in current systems, like the file size, name and permission mode. However, only a limited amount of \textit{rich metadata}, which record not only the attributes of entities, but also relationships between them, are captured in current HPC systems. The main challenge is that the rich metadata can include huge amounts of data from many sources, including users and applications, and generally must be combined to present a correct view for later query and processing. So, collecting, integrating, processing, and querying such rich metadata place a huge pressure on HPC systems. In this paper, we propose a rich metadata management approach that unifies metadata into one generic and flexible property graph. We argue that this approach both suppo
rt simple metadata operations, including directory traversal, permission validation, and, more importantly, support rich metadata processing, like provenance storage and query. The benefits of this approach come from the unified way of managing all metadata and also from the rapid evolving graph storage and processing techniques.
+HPC platforms are capable of generating huge amounts of metadata about different entities including jobs, users, and files. \textit{Simple metadata}, which describe the attributes of these entities, like the file size, name, and permissions mode, has already been well recorded and used in current systems. However, only a limited amount of \textit{rich metadata}, which record not only the attributes of entities, but also relationships between them, are captured in current HPC systems. The main challenge is that rich metadata can include a huge amount of data from many sources, including users and applications, and generally must be combined to present a useful view for later query and processing. So, collecting, integrating, processing, and querying such rich metadata place a huge pressure on HPC systems. In this paper, we propose a rich metadata management approach that unifies metadata into one generic property graph. We argue that this approach both supports simple metadat
a operations, including directory traversal, permission validation, and, more importantly, support rich metadata processing, such as provenance storage and query. The benefits of this approach come from the unified way of managing all metadata and also from leveraging rapidly evolving graph storage and processing techniques.
%For example, the rich metadata \textit{provenance}, which records the entire life cycle of data objects, is not well supported, even given the fact that provenance provides many appealing data management functionalities, such as determining the quality of the data set, finding the source of data corruption, and tracing all data dependency etc.
diff --git a/papers/meta-graph/ack.tex b/papers/meta-graph/ack.tex
index e69de29..1d8a210 100644
--- a/papers/meta-graph/ack.tex
+++ b/papers/meta-graph/ack.tex
@@ -0,0 +1,4 @@
+\section*{Acknowledgment}
+
+This material is based upon work supported by he United States Department of Defense, and the U.S. Department of Energy, Office of
+Science, under Contract No. DE-AC02-06CH11357, and by also the National Science Foundation under grant CNS-1338078 and CNS-1162488.
\ No newline at end of file
diff --git a/papers/meta-graph/bib.bib b/papers/meta-graph/bib.bib
index 79a36a5..57ed1b7 100644
--- a/papers/meta-graph/bib.bib
+++ b/papers/meta-graph/bib.bib
@@ -1,9 +1,59 @@
+ @inproceedings{facebookgs,
+ author = {Avery Ching},
+ title = {Giraph: Production-Grade Graph Processing Infrastructure for Trillion Edge Graphs},
+ booktitle = {ATPESC},
+ series = {ATPESC '14},
+ year = {2014},
+ location = {Chicago, IL},
+}
+@inproceedings{leung2009spyglass,
+ title={Spyglass: Fast, Scalable Metadata Search for Large-Scale Storage Systems.},
+ author={Leung, Andrew W and Shao, Minglong and Bisson, Timothy and Pasupathy, Shankar and Miller, Ethan L},
+ booktitle={FAST},
+ volume={9},
+ pages={153--166},
+ year={2009}
+}
+
+@article{clauset2009power,
+ title={Power-law Distributions in Empirical Data},
+ author={Clauset, Aaron and Shalizi, Cosma Rohilla and Newman, Mark EJ},
+ journal={SIAM review},
+ volume={51},
+ number={4},
+ pages={661--703},
+ year={2009},
+ publisher={SIAM}
+ }
+
@conference{provenwiki,
Booktitle = {http://en.wikipedia.org/wiki/Provenance},
Owner = {Wikipedia},
Quality = {1},
Title = {Provenance}
}
+@inproceedings{Muniswamy-Reddy:2006:PSS:1267359.1267363,
+ author = {Muniswamy-Reddy, Kiran-Kumar and Holland, David A. and Braun, Uri and Seltzer, Margo},
+ title = {Provenance-aware Storage Systems},
+ booktitle = {Proceedings of the Annual Conference on USENIX '06 Annual Technical Conference},
+ series = {ATEC '06},
+ year = {2006},
+ location = {Boston, MA},
+ pages = {4--4},
+ numpages = {1},
+ url = {http://dl.acm.org/citation.cfm?id=1267359.1267363},
+ acmid = {1267363},
+ publisher = {USENIX Association},
+ address = {Berkeley, CA, USA},
+}
+@incollection{buneman2001and,
+ title={Why and where: A Characterization of Data Provenance},
+ author={Buneman, Peter and Khanna, Sanjeev and Wang-Chiew, Tan},
+ booktitle={Database Theory ICDT 2001},
+ pages={316--330},
+ year={2001},
+ publisher={Springer}
+}
@inproceedings{yang2012defining,
title={Defining and Evaluating Network Communities based on Ground-truth},
@@ -99,8 +149,8 @@
title = {Titan},
howpublished = {\url{http://thinkaurelius.github.io/titan/}}
}
-@misc{gigraph,
- title = {Gigraph},
+@misc{giraph,
+ title = {Giraph},
howpublished = {\url{http://giraph.apache.org/}}
}
@book{berge1973graphs,
diff --git a/papers/meta-graph/conclusion.tex b/papers/meta-graph/conclusion.tex
index 82f0db2..0bad2dc 100644
--- a/papers/meta-graph/conclusion.tex
+++ b/papers/meta-graph/conclusion.tex
@@ -1,3 +1,3 @@
\section{Conclusion \& Future Work}
-In this paper, we proposed an idea of unifying rich metadata in a HPC platform into a property graph model, which supports a wide range of metadata management requirements in a simple and efficient way. By prototyping such a metadata graph from Darshan I/O traces of a real world leading supercomputer, we explorer the attributes of such graphs and compare it with some popular big graphs. After that, we introduce the existed graph facilities and discuss their feasibilities and limitations in our scenario. In general, we argue the benefits of unifying HPC metadata into a graph and also present the feasibility of implementing such a graph in current HPC platform. In future, the work will be implementing such a platform with optimized or tweaked graph facilities and provide a practical metadata solution for Exscale data management challenge.
\ No newline at end of file
+In this paper, we proposed an idea of unifying rich metadata in a HPC platform into a property graph model, which supports a wide range of metadata management functionalities in a simple but efficient way. By prototyping such a metadata graph from Darshan I/O traces of real world leading supercomputer, we explored the attributes of such graphs. Based on the prototype, we introduce existed graph facilities and discuss their remained challenges for storing and processing the metadata graph. In general, we argue the benefits of unifying HPC metadata into a graph, and also present the feasibility of implementing such a graph in current HPC platform. In the future, the work will be implementing such a platform with optimized graph facilities to provide a practical metadata solution for Exa-scale data management challenge.
\ No newline at end of file
diff --git a/papers/meta-graph/design.tex b/papers/meta-graph/design.tex
index 899c032..c8aa360 100644
--- a/papers/meta-graph/design.tex
+++ b/papers/meta-graph/design.tex
@@ -1,21 +1,21 @@
\section{Graph Facilities \& Challenges}
-Graph databases and distributed processing frameworks are two basic facilities to support our metadata graph model. Although there are already a large number of them, the challenges are still existing due to the specific requirements of storing, processing, and querying the large-scale metadata graphs.
+Graph databases and distributed processing frameworks are two basic facilities to support our metadata graph model. Although there are already a large number of them, challenges still exist due to the specific requirements of storing, processing, and querying the metadata graphs of this scale.
\subsection{Graph Databases}
-The graph databases are designed to cover the requirements of complex graph-based relationships, which embarrass the traditional relational databases. In the last couple of years, there have been an increasing number of graph databases implementation, including AllegroGraph, DEX, G-Store, HyperGraphDB, InfinitGraphDB, Neo4j, and Titan etc~\cite{allegrograph, dex, steinhaus2010g, iordanov2010hypergraphdb, igraph, webber2012programmatic, titan}. They can be categorized based on different metrics. Based on storage device, there are in-memory databases and disk-based databases. Based on the supported graph data structure, there are {simple graphs} databases, {hypergraphs} databases, and {property graphs} databases\footnote{Here, the simple graph indicates graph defined as a set of nodes connected by weighted edges. Hypergraphs extends the simple graphs by allowing an edge to relate an arbitrary number of nodes~\cite{berge1973graphs}. Property graph indicates graph where nodes a
nd edges contain properties. This property graph is the very basis of our proposed metadata graph model.}. Based on distributed deployment, there are single server databases, high availability databases, and distributed databases.\footnote{Here, the difference between high availability (HA) and distribution is HA only provides distributed reads on identical copies of the same dataset.}
+The graph databases are designed to cover the requirements of complex graph-based relationships, which are not well-suited to traditional relational databases. In last couple of years, there have been an increasing number of graph database implementations, including AllegroGraph, DEX, G-Store, HyperGraphDB, InfinitGraphDB, Neo4j, and Titan~\cite{allegrograph, dex, steinhaus2010g, iordanov2010hypergraphdb, igraph, webber2012programmatic, titan}. They can be categorized based on different metrics. Based on storage device, there are in-memory databases and disk-based databases. Based on the supported graph data structure, there are {simple graphs} databases, {hypergraphs} databases, and {property graphs} databases\footnote{Here, the simple graph indicates graph defined as a set of nodes connected by weighted edges. Hypergraphs extends the simple graphs by allowing an edge to relate an arbitrary number of nodes~\cite{berge1973graphs}. Property graph indicates graph where nodes
and edges contain properties. This property graph is the very basis of our proposed metadata graph model.}. Based on distributed deployment, there are single server databases, high availability databases, and distributed databases.\footnote{Here, the difference between high availability (HA) and distribution is HA only provides distributed reads on identical copies of the same dataset.}
%AllegroGraph and G-Store belong to this category.
%Databases like HyperGraphDB and Sones support these hypergraphs.
%Many databases like Neo4j, Titan, InfinitGraph, and DEX support such graphs.
-Among those graph databases, metadata graphs first requires a disk-based solution as they are too large to fit into memory; second, the supported graph structures should be property graph model. Also, the distribution supports will also be necessary since the graph size may overflow the disks of a single sever. There are several implementations that satisfy such requirements, like Titan, DEX, Neo4j etc. But, the challenge is the performance. Updating edges across different servers will significantly reduce the performance, so the graph databases should consider the structure of different entities and provide intelligent storage layout. Another key performance challenge is graph traversal. In fact, traveling through a property graph like our metadata graph will be even more complex. The main reason is that, during traveling, we usually need to apply filter or computations on some properties like the provenance example shows. As these properties are too big to be fully cached
in memory and loading them from persistent devices will introduce too many random seeking, we will need a intelligent cache strategy in our design too.
+Among those graph databases, first, the metadata graphs are based on the property graph model; second, they require disk-based solutions as they are too large to fit into memory. Also, the distribution supports will also be necessary since the graph size may overflow the disks of a single sever. There are several implementations that satisfy such requirements, such as Titan, DEX, and Neo4j. But, the challenge is performance. First, in distributed environment, any updating across different servers will significantly reduce the performance, so we need to consider the structure of the metadata graph and provide optimized storage layout. Another performance challenge is graph traversal. In fact, traveling through metadata graph usually includes applying filters and computations on properties during traveling just like the provenance example in previous section shows. As these properties are too big to be fully cached in memory, we have to load them from persistent devices each
time. This will introduce too many random seeks, it is clear that we will need an intelligent cache strategy here.
%the graph size, graph data structure, and storage layouts, all determine the traversal performance. For a moderate-sized simple graph. we may be able to perform traversal in memory. However,
\subsection{Graph Processing}
-In addition to the graph databases, there are also graph processing frameworks which can be used to perform computation or queries on graphs in a distributed way. Typical examples of these frameworks include Gigraph~\cite{gigraph}, which was designed and implemented based on Pregel computing model~\cite{malewicz2010pregel}; GraphX~\cite{xin2013graphx}, which was based on Spark computing framework~\cite{zaharia2010spark}; GraphLab~\cite{low2010graphlab}, and X-Stream~\cite{roy2013x} etc. Those processing frameworks are a complement of the querying and searching from graph databases. For example, we can run \textit{community discovery} algorithms on metadata graph to find the `closely' data files. The results can be used to optimize physical placements for better I/O performance. These algorithms usually get the whole graph involved in the iterative computation and last a long time.
+In addition to the graph databases, there are also graph processing frameworks that can be used to perform computation or queries on graphs in a distributed way. Typical examples of these frameworks include Giraph~\cite{giraph}, which was designed and implemented based on Pregel computing model~\cite{malewicz2010pregel}; GraphX~\cite{xin2013graphx}, which was based on Spark computing framework~\cite{zaharia2010spark}; GraphLab~\cite{low2010graphlab}, and X-Stream~\cite{roy2013x}. Those processing frameworks are a complement of the querying and searching from graph databases. For example, we can run \textit{community discovery} algorithms on metadata graph to find the `closely' related data files. The results can be used to optimize physical placements for better I/O performance. These algorithms usually get the whole graph involved into an iterative computation that lasts a long time.
%CSR or CSC
-However, most these distributed graph processing frameworks work on the unstructured graphs which are usually simply stored as a plain file in adjacency list formats in a general storage back-end. So, there is a big gap from deploying graph algorithms on these plain graph formats to running these algorithms on a graph database, which has an optimized storage layout for querying and searching.
+However, most these distributed graph processing frameworks were designed to work on the unstructured graphs which are usually simply stored as a plain file in adjacency list formats in a general storage back-end. There is a big gap from deploying graph algorithms in these plain graph formats to running these algorithms on a graph database, which usually has an specific, optimized storage layout for querying and searching.
%However, the ability to run complex graph algorithms is necessary for metadata management, and this ability can not be easily satisfied using the querying or searching facilities provided by graph databases. A distributed fault-tolerant graph processing model is a much better choice than writing applications to manage all these complexity.
diff --git a/papers/meta-graph/exps/2013-graph.numbers/Index.zip b/papers/meta-graph/exps/2013-graph.numbers/Index.zip
index 52002d3..ab2ff50 100644
Binary files a/papers/meta-graph/exps/2013-graph.numbers/Index.zip and b/papers/meta-graph/exps/2013-graph.numbers/Index.zip differ
diff --git a/papers/meta-graph/exps/2013-graph.numbers/Metadata/Properties.plist b/papers/meta-graph/exps/2013-graph.numbers/Metadata/Properties.plist
index b6bd2e9..4ae97f2 100644
Binary files a/papers/meta-graph/exps/2013-graph.numbers/Metadata/Properties.plist and b/papers/meta-graph/exps/2013-graph.numbers/Metadata/Properties.plist differ
diff --git a/papers/meta-graph/exps/fdegree.pdf b/papers/meta-graph/exps/fdegree.pdf
index 4ca0abf..2911b4b 100644
Binary files a/papers/meta-graph/exps/fdegree.pdf and b/papers/meta-graph/exps/fdegree.pdf differ
diff --git a/papers/meta-graph/exps/jdegree.pdf b/papers/meta-graph/exps/jdegree.pdf
index 2e379e7..2408949 100644
Binary files a/papers/meta-graph/exps/jdegree.pdf and b/papers/meta-graph/exps/jdegree.pdf differ
diff --git a/papers/meta-graph/exps/pdegree.pdf b/papers/meta-graph/exps/pdegree.pdf
index a2b5aec..06457b5 100644
Binary files a/papers/meta-graph/exps/pdegree.pdf and b/papers/meta-graph/exps/pdegree.pdf differ
diff --git a/papers/meta-graph/exps/udegree.pdf b/papers/meta-graph/exps/udegree.pdf
index 127a3c0..b7515c7 100644
Binary files a/papers/meta-graph/exps/udegree.pdf and b/papers/meta-graph/exps/udegree.pdf differ
diff --git a/papers/meta-graph/intro.tex b/papers/meta-graph/intro.tex
index 05c1dbf..d282136 100644
--- a/papers/meta-graph/intro.tex
+++ b/papers/meta-graph/intro.tex
@@ -1,21 +1,21 @@
\section{Introduction}
-Metadata, especially rich metadata, contains detailed information about different entities and their relationships. These entities could be users, jobs, processes, data files, or even user-defined entities. Storing and utilizing metadata already provides the basic data management functionalities in existing storage systems, including finding files, controlling file access, and tracing file creation and access time. We categorize this metadata as \textit{simple metadata} since they only contain the predefined attributes about individual entities and very basic relationships (e.g. ownership, POSIX namespace). On the contrary, \textit{rich metadata} include more than individual predefined attributes; they may store user-defined arbitrary attributes of entities and even their relationships. A typical example of rich metadata would be provenance~\cite{provenwiki} (e.g. lineage).
+Metadata, especially rich metadata, contains detailed information about different entities and their relationships. These entities could be users, jobs, processes, data files, or other user-defined entities. Storing and utilizing metadata already provide basic data management functionalities in existing storage systems, including locating files, controlling file access, and tracing file creation time. Systems like Spyglass~\cite{leung2009spyglass} were also designed to provide enhanced searching functionality to utilize those metadata. Generally, we categorize these metadata as \textit{simple metadata} since they only contain the predefined attributes about individual entities and very basic relationships (e.g. ownership, POSIX namespace). On the contrary, \textit{rich metadata} include more than individual predefined attributes: they store user-defined arbitrary attributes of entities and their relationships. A typical example of rich metadata would be provenance~\cite{bune
man2001and, Muniswamy-Reddy:2006:PSS:1267359.1267363} (e.g. lineage).
-Provenance is well understood in the context of art or digital libraries, where it respectively refers to the documented history of an art object, or the documentation of processes in a digital object's life cycle~\cite{moreau2011open}. In the computational systems, it indicates a recording of complete history of each data element, including the processes that generated it, the user that started the processes, and even the environment variables, parameters, and configuration files while executing. A complete provenance picture supports a huge amount of data management abilities~\cite{simmhan2005survey, silva2007provenance}. For example, the accessing history of users reading/writing data files can help us develop an audit tool to monitor and administrate users in shared supercomputer facilities; the detailed read/write history from processes to data pieces provides a possibility to trace back suspicious executions that generated or were based on wrong datasets; reproducibil
ity also may be possible because we have the complete history of an execution and have a better chance to re-generate the same environment to run it again.
+Provenance is well understood in the context of art or digital libraries, where it respectively refers to the documented history of an art object, or the documentation of processes in a digital object's life cycle~\cite{moreau2011open}. In the computational systems, it indicates a recording of complete history of each data element, including the processes that generated it, the user that started the processes, and even the environment variables, parameters, and configuration files while executing. A complete provenance picture supports a huge amount of data management abilities~\cite{simmhan2005survey, silva2007provenance}. For example, the access history of users reading/writing data files can help monitor users in shared supercomputer facilities; the detailed read/write history from processes to data provides a possibility to trace back suspicious executions; reproducibility also may be possible because we have the complete history of an execution to re-generate the same
environment to reproduce the results.
-While there are numerous advantages to capture rich metadata like provenance, current HPC platforms still lack basic facilities to collect, store, process, and query rich metadata. The challenge comes from at least three places.
+While there are numerous advantages to capturing rich metadata like provenance, current HPC platforms still lack basic facilities to collect, store, process, and query these rich metadata. There are at least three challenges. %The challenge comes from at least three places.
\begin{itemize}
-\item \textit{Storage System Pressure}. Considering a leadership supercomputer, there might be millions of processes running on millions of cores accessing billions of files per second. In this case, recording the rich metadata, like the detailed access history of each process, will place pressure on the storage system. In addition, as storing rich metadata must not affect the application execution speed significantly, the resources (both network and disks bandwidth) dedicated to storing metadata must be limited in most cases.
+\item \textit{Storage System Pressure}. Considering a leadership supercomputer, there might be millions of processes running on millions of cores accessing billions of files per second. In this case, recording rich metadata, like the detailed access history of each process, will place extreme pressure on the storage system. As storing rich metadata must not affect the application execution speed significantly, the resources (both network and disk bandwidth) dedicated to storing metadata must be limited in most cases.
-\item \textit{Efficient Processing}. Even if we can collect and store these rich metadata, it is still a big challenge to process and query them. First, as rich metadata are large and can not be held in one server, a distributed processing framework is necessary in most cases. Second, some use cases require fast searching and reading, however, some require complex queries rather than simple searching, so efficient and flexible processing should be provided for them.
+\item \textit{Efficient Processing}. Even if we can collect and store these rich metadata, it is still a big challenge to process and query them. First, as rich metadata are large and can not be held in one server, distributed processing is necessary in most cases. Second, some use cases require fast searching and reading; however, some other cases require complex queries, so efficient and flexible processing also are required for these rich metadata.
-\item \textit{Metadata Integration}. As we have described, the rich metadata could be as diverse as the users need. They can contain predefined attributes and relationships of entities, or be extended to any user-defined attributes and relationships. Traditionally, we need different tools to process them. However, this strategy leads to waste of resources as they are not able to reuse the same metadata and processing infrastructure for different use cases. %For example, the data audit application and data verification application both need to know the file access history. We either need to store this metadata in both two applications, or only store it in one application and issue lots of cross-reference reads from other application later.
+\item \textit{Metadata Integration}. As we have described, the rich metadata could be very diverse. They can contain predefined attributes and relationships of entities and be extended to any user-defined attributes and relationships. Traditionally, we use different tools to process these various data. However, this strategy leads to waste of resources as they are not able to reuse the same metadata and processing infrastructure for different use cases. %For example, the data audit application and data verification application both need to know the file access history. We either need to store this metadata in both two applications, or only store it in one application and issue lots of cross-reference reads from other application later.
\end{itemize}
%by providing an unified graph abstraction for all the entities and relationships
-In this paper, we proposed unifying all metadata into one property graph that integrates rich metadata from different sources together: applications can store their rich metadata using graph storage APIs, and access different categories of metadata using graph query APIs~\cite{jouili2013empirical}. The benefits are twofold: first, it directly solves the integration issue by using a single representation. All applications will have the same interface to store or process metadata in a single service, where we can apply complex optimizations to improve the performance further. Second, by abstracting metadata into a graph, we are able to utilize rapidly evolving graph techniques to provide better access speed, flexible query languages, and also a high-performance graph-based distributed framework.
+In this paper, we propose unifying all metadata into one property graph that integrates rich metadata from different sources: storage and resource management services and applications can store their rich metadata using graph storage APIs and access different categories of metadata using graph query APIs~\cite{jouili2013empirical}. The benefits are twofold: first, this directly solves the integration issue by using a single representation: all services and applications will use the same interface to store or process metadata in a single service where we can apply complex optimizations to improve the performance further. Second, by abstracting metadata into a graph, we are able to utilize rapidly evolving graph techniques to provide better access speed, flexible query languages, and also a high-performance graph-based distributed framework.
-This paper is organized as follows. We first introduce the definition of proposed graph model for rich metadata in Section II. In Section III, we explore the basic attributes of such graph by building an example graph using the metadata collected from the Darshan trace of a leadership supercomputer (Intrepid). In Section IV, we show the strategy to implement several critical data management functionalities based on the graph model. In Section V, we briefly introduce the relevant techniques from the graph storage and processing community, and discuss the challenges on current graph infrastructure. The last section (Section VI) concludes this paper and proposes the future work. %possible improvements
\ No newline at end of file
+This paper is organized as follows. We first introduce the proposed graph model for rich metadata in Section II. Then, we explore the attributes of such graph by building an example graph using the metadata collected from the Darshan trace of a leadership supercomputer (Intrepid) in Section III. In Section IV, we show the strategy to implement several critical data management functionalities based on the graph model. In Section V, we briefly introduce the relevant techniques from the graph storage and processing community, and discuss the challenges on current graph infrastructure. The last section (Section VI) concludes this paper and proposes the future work. %possible improvements
\ No newline at end of file
diff --git a/papers/meta-graph/main.bbl b/papers/meta-graph/main.bbl
new file mode 100644
index 0000000..10729e1
--- /dev/null
+++ b/papers/meta-graph/main.bbl
@@ -0,0 +1,222 @@
+% Generated by IEEEtran.bst, version: 1.13 (2008/09/30)
+\begin{thebibliography}{10}
+\providecommand{\url}[1]{#1}
+\csname url@samestyle\endcsname
+\providecommand{\newblock}{\relax}
+\providecommand{\bibinfo}[2]{#2}
+\providecommand{\BIBentrySTDinterwordspacing}{\spaceskip=0pt\relax}
+\providecommand{\BIBentryALTinterwordstretchfactor}{4}
+\providecommand{\BIBentryALTinterwordspacing}{\spaceskip=\fontdimen2\font plus
+\BIBentryALTinterwordstretchfactor\fontdimen3\font minus
+ \fontdimen4\font\relax}
+\providecommand{\BIBforeignlanguage}[2]{{%
+\expandafter\ifx\csname l@#1\endcsname\relax
+\typeout{** WARNING: IEEEtran.bst: No hyphenation pattern has been}%
+\typeout{** loaded for the language `#1'. Using the pattern for}%
+\typeout{** the default language instead.}%
+\else
+\language=\csname l@#1\endcsname
+\fi
+#2}}
+\providecommand{\BIBdecl}{\relax}
+\BIBdecl
+
+\bibitem{leung2009spyglass}
+A.~W. Leung, M.~Shao, T.~Bisson, S.~Pasupathy, and E.~L. Miller, ``Spyglass:
+ Fast, Scalable Metadata Search for Large-Scale Storage Systems.'' in
+ \emph{FAST}, vol.~9, 2009, pp. 153--166.
+
+\bibitem{buneman2001and}
+P.~Buneman, S.~Khanna, and T.~Wang-Chiew, ``Why and where: A Characterization
+ of Data Provenance,'' in \emph{Database Theory ICDT 2001}.\hskip 1em plus
+ 0.5em minus 0.4em\relax Springer, 2001, pp. 316--330.
+
+\bibitem{Muniswamy-Reddy:2006:PSS:1267359.1267363}
+\BIBentryALTinterwordspacing
+K.-K. Muniswamy-Reddy, D.~A. Holland, U.~Braun, and M.~Seltzer,
+ ``Provenance-aware Storage Systems,'' in \emph{Proceedings of the Annual
+ Conference on USENIX '06 Annual Technical Conference}, ser. ATEC '06.\hskip
+ 1em plus 0.5em minus 0.4em\relax Berkeley, CA, USA: USENIX Association, 2006,
+ pp. 4--4. [Online]. Available:
+ \url{http://dl.acm.org/citation.cfm?id=1267359.1267363}
+\BIBentrySTDinterwordspacing
+
+\bibitem{moreau2011open}
+L.~Moreau, B.~Clifford, J.~Freire, J.~Futrelle, Y.~Gil, P.~Groth,
+ N.~Kwasnikowska, S.~Miles, P.~Missier, J.~Myers \emph{et~al.}, ``The Open
+ Provenance Model Core Specification (v1. 1),'' \emph{Future Generation
+ Computer Systems}, vol.~27, no.~6, pp. 743--756, 2011.
+
+\bibitem{simmhan2005survey}
+Y.~L. Simmhan, B.~Plale, and D.~Gannon, ``A Survey of Data Provenance in
+ e-science,'' \emph{ACM Sigmod Record}, vol.~34, no.~3, pp. 31--36, 2005.
+
+\bibitem{silva2007provenance}
+C.~T. Silva, J.~Freire, and S.~P. Callahan, ``Provenance for Visualizations:
+ Reproducibility and Beyond,'' \emph{Computing in Science \& Engineering},
+ vol.~9, no.~5, pp. 82--89, 2007.
+
+\bibitem{jouili2013empirical}
+S.~Jouili and V.~Vansteenberghe, ``An Empirical Comparison of Graph
+ Databases,'' in \emph{Social Computing (SocialCom), 2013 International
+ Conference on}.\hskip 1em plus 0.5em minus 0.4em\relax IEEE, 2013, pp.
+ 708--715.
+
+\bibitem{tanenbaum1992modern}
+A.~S. Tanenbaum and A.~Tannenbaum, \emph{Modern Operating Systems}.\hskip 1em
+ plus 0.5em minus 0.4em\relax Prentice hall Englewood Cliffs, 1992, vol.~2.
+
+\bibitem{propertygraph}
+``Property Graph,'' \url{http://www.w3.org/community/propertygraphs/}.
+
+\bibitem{muniswamy2006provenance}
+K.-K. Muniswamy-Reddy, D.~A. Holland, U.~Braun, and M.~I. Seltzer,
+ ``Provenance-Aware Storage Systems.'' in \emph{USENIX Annual Technical
+ Conference, General Track}, 2006, pp. 43--56.
+
+\bibitem{muniswamy2009layering}
+K.-K. Muniswamy-Reddy, U.~Braun, D.~A. Holland, P.~Macko, D.~Maclean, D.~Margo,
+ M.~Seltzer, and R.~Smogor, ``Layering in Provenance Systems,'' in
+ \emph{Proceedings of the 2009 USENIX Annual Technical Conference}, 2009.
+
+\bibitem{braun2006issues}
+U.~Braun, S.~Garfinkel, D.~A. Holland, K.-K. Muniswamy-Reddy, and M.~I.
+ Seltzer, ``Issues in Automatic Provenance Collection,'' in \emph{Provenance
+ and annotation of data}.\hskip 1em plus 0.5em minus 0.4em\relax Springer,
+ 2006, pp. 171--183.
+
+\bibitem{carns200924}
+P.~Carns, R.~Latham, R.~Ross, K.~Iskra, S.~Lang, and K.~Riley, ``24/7
+ Characterization of Petascale I/O Workloads,'' in \emph{Cluster Computing and
+ Workshops, 2009. CLUSTER'09. IEEE International Conference on}.\hskip 1em
+ plus 0.5em minus 0.4em\relax IEEE, 2009, pp. 1--10.
+
+\bibitem{carns2011understanding}
+P.~Carns, K.~Harms, W.~Allcock, C.~Bacon, S.~Lang, R.~Latham, and R.~Ross,
+ ``Understanding and Improving Computational Science Storage Access through
+ Continuous Characterization,'' \emph{ACM Transactions on Storage (TOS)},
+ vol.~7, no.~3, p.~8, 2011.
+
+\bibitem{darshanlog2013}
+``FTP site: Darshan data.'' \url{ftp://ftp.mcs.anl.gov/pub/darshan/data/}.
+
+\bibitem{demetrescu2009shortest}
+C.~Demetrescu, A.~V. Goldberg, and D.~S. Johnson, \emph{The Shortest Path
+ Problem: Ninth DIMACS Implementation Challenge}.\hskip 1em plus 0.5em minus
+ 0.4em\relax American Mathematical Soc., 2009, vol.~74.
+
+\bibitem{facebookgs}
+A.~Ching, ``Giraph: Production-Grade Graph Processing Infrastructure for
+ Trillion Edge Graphs,'' in \emph{ATPESC}, ser. ATPESC '14, 2014.
+
+\bibitem{guillaume2002web}
+J.-L. Guillaume, M.~Latapy \emph{et~al.}, ``The Web Graph: an Overview,'' in
+ \emph{Actes d'ALGOTEL'02 (Quatri{\`e}mes Rencontres Francophones sur les
+ aspects Algorithmiques des T{\'e}l{\'e}communications)}, 2002.
+
+\bibitem{twitters}
+``Twitter Statistics,''
+ \url{http://www.statisticbrain.com/twitter-statistics/}.
+
+\bibitem{twitteredge}
+``Twitter Edge Statistics,''
+ \url{http://www.theguardian.com/technology/blog/2009/jun/29/twitter-users-average-api-traffic}.
+
+\bibitem{faloutsos1999power}
+M.~Faloutsos, P.~Faloutsos, and C.~Faloutsos, ``On Power-Law Relationships of
+ the Internet Topology,'' in \emph{ACM SIGCOMM Computer Communication Review},
+ vol.~29, no.~4.\hskip 1em plus 0.5em minus 0.4em\relax ACM, 1999, pp.
+ 251--262.
+
+\bibitem{clauset2009power}
+A.~Clauset, C.~R. Shalizi, and M.~E. Newman, ``Power-law Distributions in
+ Empirical Data,'' \emph{SIAM review}, vol.~51, no.~4, pp. 661--703, 2009.
+
+\bibitem{patil2011scale}
+S.~Patil and G.~A. Gibson, ``Scale and Concurrency of GIGA+: File System
+ Directories with Millions of Files.'' in \emph{FAST}, vol.~11, 2011, pp.
+ 13--13.
+
+\bibitem{kim2012sbv}
+M.~Kim and K.~S. Candan, ``SBV-Cut: Vertex-cut based Graph Partitioning Using
+ Structural Balance Vertices,'' \emph{Data \& Knowledge Engineering}, vol.~72,
+ pp. 285--303, 2012.
+
+\bibitem{gonzalez2012powergraph}
+J.~E. Gonzalez, Y.~Low, H.~Gu, D.~Bickson, and C.~Guestrin, ``PowerGraph:
+ Distributed Graph-Parallel Computation on Natural Graphs.'' in \emph{OSDI},
+ vol.~12, no.~1, 2012, p.~2.
+
+\bibitem{abou2006multilevel}
+A.~Abou-Rjeili and G.~Karypis, ``Multilevel Algorithms for Partitioning
+ Power-Law Graphs,'' in \emph{Parallel and Distributed Processing Symposium,
+ 2006. IPDPS 2006. 20th International}.\hskip 1em plus 0.5em minus 0.4em\relax
+ IEEE, 2006, pp. 10--pp.
+
+\bibitem{provchallengeweb}
+``Provenance Challenges,'' \url{http://twiki.ipaw.info/bin/view/Challenge/}.
+
+\bibitem{allegrograph}
+``AllegroGraph,'' \url{http://franz.com/agraph/allegrograph/}.
+
+\bibitem{dex}
+``DEX,'' \url{http://www.sparsity-technologies.com/}.
+
+\bibitem{steinhaus2010g}
+R.~Steinhaus, D.~Olteanu, and T.~Furche, ``G-Store: A Storage Manager for Graph
+ Data,'' Ph.D. dissertation, Citeseer, 2010.
+
+\bibitem{iordanov2010hypergraphdb}
+B.~Iordanov, ``HyperGraphDB: a Generalized Graph Database,'' in \emph{Web-Age
+ Information Management}.\hskip 1em plus 0.5em minus 0.4em\relax Springer,
+ 2010, pp. 25--36.
+
+\bibitem{igraph}
+``InfiniteGraph,'' \url{http://www.objectivity.com/infinitegraph}.
+
+\bibitem{webber2012programmatic}
+J.~Webber, ``A Programmatic Introduction to Neo4j,'' in \emph{Proceedings of
+ the 3rd annual conference on Systems, programming, and applications: software
+ for humanity}.\hskip 1em plus 0.5em minus 0.4em\relax ACM, 2012, pp.
+ 217--218.
+
+\bibitem{titan}
+``Titan,'' \url{http://thinkaurelius.github.io/titan/}.
+
+\bibitem{berge1973graphs}
+C.~Berge and E.~Minieka, \emph{Graphs and Hypergraphs}.\hskip 1em plus 0.5em
+ minus 0.4em\relax North-Holland publishing company Amsterdam, 1973, vol.~7.
+
+\bibitem{giraph}
+``Giraph,'' \url{http://giraph.apache.org/}.
+
+\bibitem{malewicz2010pregel}
+G.~Malewicz, M.~H. Austern, A.~J. Bik, J.~C. Dehnert, I.~Horn, N.~Leiser, and
+ G.~Czajkowski, ``Pregel: a System for Large-Scale Graph Processing,'' in
+ \emph{Proceedings of the 2010 ACM SIGMOD International Conference on
+ Management of data}.\hskip 1em plus 0.5em minus 0.4em\relax ACM, 2010, pp.
+ 135--146.
+
+\bibitem{xin2013graphx}
+R.~S. Xin, J.~E. Gonzalez, M.~J. Franklin, and I.~Stoica, ``GraphX: A Resilient
+ Distributed Graph System on Spark,'' in \emph{First International Workshop on
+ Graph Data Management Experiences and Systems}.\hskip 1em plus 0.5em minus
+ 0.4em\relax ACM, 2013, p.~2.
+
+\bibitem{zaharia2010spark}
+M.~Zaharia, M.~Chowdhury, M.~J. Franklin, S.~Shenker, and I.~Stoica, ``Spark:
+ Cluster Computing with Working Sets,'' in \emph{Proceedings of the 2nd USENIX
+ conference on Hot topics in cloud computing}, 2010, pp. 10--10.
+
+\bibitem{low2010graphlab}
+Y.~Low, J.~Gonzalez, A.~Kyrola, D.~Bickson, C.~Guestrin, and J.~M. Hellerstein,
+ ``Graphlab: A New Framework for Parallel Machine Learning,'' \emph{arXiv
+ preprint arXiv:1006.4990}, 2010.
+
+\bibitem{roy2013x}
+A.~Roy, I.~Mihailovic, and W.~Zwaenepoel, ``X-stream: Edge-Centric Graph
+ Processing using Streaming Partitions,'' in \emph{Proceedings of the
+ Twenty-Fourth ACM Symposium on Operating Systems Principles}.\hskip 1em plus
+ 0.5em minus 0.4em\relax ACM, 2013, pp. 472--488.
+
+\end{thebibliography}
diff --git a/papers/meta-graph/main.blg b/papers/meta-graph/main.blg
new file mode 100644
index 0000000..3f8f4a0
--- /dev/null
+++ b/papers/meta-graph/main.blg
@@ -0,0 +1,56 @@
+This is BibTeX, Version 0.99d (TeX Live 2013)
+Capacity: max_strings=35307, hash_size=35307, hash_prime=30011
+The top-level auxiliary file: main.aux
+The style file: IEEEtran.bst
+Reallocated singl_function (elt_size=4) to 100 items from 50.
+Reallocated singl_function (elt_size=4) to 100 items from 50.
+Reallocated singl_function (elt_size=4) to 100 items from 50.
+Reallocated wiz_functions (elt_size=4) to 6000 items from 3000.
+Reallocated singl_function (elt_size=4) to 100 items from 50.
+Database file #1: bib.bib
+-- IEEEtran.bst version 1.13 (2008/09/30) by Michael Shell.
+-- http://www.michaelshell.org/tex/ieeetran/bibtex/
+-- See the "IEEEtran_bst_HOWTO.pdf" manual for usage information.
+
+Done.
+You've used 41 entries,
+ 4035 wiz_defined-function locations,
+ 1017 strings with 14069 characters,
+and the built_in function-call counts, 26767 in all, are:
+= -- 2017
+> -- 651
+< -- 182
++ -- 345
+- -- 119
+* -- 1225
+:= -- 3920
+add.period$ -- 94
+call.type$ -- 41
+change.case$ -- 0
+chr.to.int$ -- 417
+cite$ -- 41
+duplicate$ -- 2019
+empty$ -- 2469
+format.name$ -- 147
+if$ -- 6223
+int.to.chr$ -- 0
+int.to.str$ -- 41
+missing$ -- 387
+newline$ -- 148
+num.names$ -- 31
+pop$ -- 997
+preamble$ -- 1
+purify$ -- 0
+quote$ -- 2
+skip$ -- 2122
+stack$ -- 0
+substring$ -- 1016
+swap$ -- 1490
+text.length$ -- 43
+text.prefix$ -- 0
+top$ -- 5
+type$ -- 41
+warning$ -- 0
+while$ -- 114
+width$ -- 43
+write$ -- 376
diff --git a/papers/meta-graph/main.tex b/papers/meta-graph/main.tex
index 3ee8b25..b94dfe2 100644
--- a/papers/meta-graph/main.tex
+++ b/papers/meta-graph/main.tex
@@ -2,7 +2,7 @@
\usepackage[numbers]{natbib}
%\usepackage{subfig}
\usepackage[subfigure]{graphfig}
-%\usepackage{flushend}
+\usepackage{flushend}
%\usepackage{cite}
%\usepackage{enumitem}
\usepackage{color}
@@ -16,9 +16,8 @@
\usepackage{listings}
\usepackage[noblocks]{authblk}
\usepackage{epstopdf}
-
+\usepackage{draftwatermark}
\usepackage{url}
-
\definecolor{Gray}{gray}{0.9}
\hyphenation{op-tical net-works semi-conduc-tor}
\algrenewcommand{\algorithmiccomment}[1]{\hskip3em$\rightarrow$ #1}
@@ -69,15 +68,15 @@
xrightmargin=0.5em,
}
\begin{document}
-\title{On the Use of Property Graph for Rich Metadata Management in HPC Systems}
+\title{Using Property Graphs for Rich Metadata Management in HPC Systems}
%\author[1]{Dong Dai}
+%\author[2]{Robert Ross}
%\author[1]{Yong Chen}
%\author[2]{Phil Carns}
-%\author[2]{Robert Ross}
%\author[2]{Dries Kimpe}
-%\affil[1]{Computer Science Department, Texas Tech University, USA, \{dong.dai, yong.chen\}(a)ttu.edu}
-%\affil[2]{Mathematics and Computer Science Division, Argonne National Laboratory, USA, \{dkimpe, pcarns, rross\}(a)mcs.anl.gov}
+%\affil[1]{Computer Science Department, Texas Tech University, USA, \{dong.dai, %yong.chen\}(a)ttu.edu}
+%\affil[2]{Mathematics and Computer Science Division, Argonne National Laboratory, USA, %\{dkimpe, pcarns, rross\}(a)mcs.anl.gov}
\maketitle
\input{abstract}
diff --git a/papers/meta-graph/model.tex b/papers/meta-graph/model.tex
index d928853..fdf06bd 100644
--- a/papers/meta-graph/model.tex
+++ b/papers/meta-graph/model.tex
@@ -1,26 +1,26 @@
\section{Graph-based Metadata Model}
%ritchie1978unix
-In fact, we already consider metadata as a graph. The traditional directory-based file management constructs a tree structure to manage files with additional metadata stored in \textit{inodes} at leaves in the tree~\cite{tanenbaum1992modern}. This tree is a graph. The provenance standard (\textit{Open Provenance Model}~\cite{moreau2011open}) considers the provenance of objects is represented by an annotated causality graph, which is a directed acyclic graph enriched with annotations capturing further information.
+In fact, we already consider metadata as a graph. The traditional directory-based file management constructs a tree structure to manage files with additional metadata stored in \textit{inodes}~\cite{tanenbaum1992modern}. This tree is a graph. The provenance standard (\textit{Open Provenance Model}~\cite{moreau2011open}) captures the provenance of objects by an annotated causality graph, which is a directed acyclic graph enriched with annotations capturing further information.
-We generalize these graphs in HPC scenarios and propose the metadata graph model. The metadata graph is derived from the \textit{property graph model}~\cite{propertygraph}, which includes vertice that represent entities in the system, edges that show their relationships, and properties that annotate both vertice and edges and can store arbitrary information users want. Based on the entities in HPC environment, we introduce the strategy to map the possibly arbitrary rich metadata into this property graph model.
+We generalize these graphs in HPC scenarios and propose the metadata graph model. The metadata graph is derived from the \textit{property graph model}~\cite{propertygraph}, which includes vertices that represent entities in the system, edges that show their relationships, and properties that annotate both vertice and edges and can store arbitrary information users want. Based on the entities in an HPC environment, we introduce the strategy to map the possibly arbitrary rich metadata into this property graph model.
\subsection{Entity To Vertex}
-In an HPC platform, there are three basic entities: users, applications, and data files. Moreover, users also can define other logical entities, like \textit{user groups} or \textit{work-flow} as they need. So, in our strategy, we define three basic entities and allow users to extend them to build more entities.
+In an HPC platform, there are three basic entities: users, the running applications, and the data files. So, we define them as three basic types of vertices as follow.
\begin{itemize}
-\item \textit{Data Object}: It represents the smallest data unit in storage systems. Each file in PFS (Parallel File System) indicates one data object. Moreover, the directory is also a data object, which contains multiple files. %The applications or users programs are also data objects.
+\item \textit{Data Object}: It represents the basic data unit in storage systems, like, in PFS (Parallel File System), each file indicates a data object. Moreover, the directory is also a data object, which may contain multiple files. %The applications or users programs are also data objects.
-\item \textit{Executions}: They represents the execution of applications. There are three levels of executions: \textit{Job} submitted by
-the user; \textit{Processes} scheduled from one job; and \textit{Threads} running inside one process. Different use cases require certain level of execution details, and generate graphs with different size. For simplicity, we name all these entities as \textit{Execution} entity in later discussion.
+\item \textit{Executions}: They represent the execution of applications. There are three levels of executions: \textit{Job} submitted by
+the user; \textit{Processes} scheduled from one job; and \textit{Threads} running inside one process. Different use cases require different levels of execution detail and generate graphs with different size. For simplicity, we name all these entities as \textit{Execution} entities in later discussion.
-\item \textit{User}: It simply means the real users of the cluster.
+\item \textit{User}: It simply means the real user of the cluster.
\end{itemize}
-In addition to these basic entities, users usually define their own entities. The user-defined entities must connect with existing entities to keep every element in the graph accessible by traveling through the graph.
+In addition to these basic entities, users can define their own entities. For example, in a work-flow system, users can create \textit{work-flow} entities and connect them with \textit{Executions} to trace work-flow executions. Also, the system administrators can create \textit{user group} entities to include different \textit{Users} and assign privileges for them. The metadata graph allows users to extend existed entities to build their own entities. The only limitation is that these user-defined entities must connect with existing entities to keep every element in the graph accessible by traveling through the graph.
-\subsection{Relation To Edge}
-There are several basic relationships between the basic entities in Table~\ref{rel}. Each cell shows the basic relationships from the row identifier to the column identifier. Each relationship denotes a directed edge in the metadata graph. For example, \textit{run} indicates that the user starts an execution; \textit{exe} means one execution is based on the data objects as the executable files; \textit{read/write} indicates the I/O operations from executions to data objects. For all those relationships, we also define the reversed ones to accelerate the reversed traversal.
+\subsection{Relationship To Edge}
+Based on the basic entities defined ahead, we have several basic relationships between them shown in Table~\ref{rel}. Each cell shows relationships from the row identifier to the column identifier. Each relationship will be mapped to a directed edge in the metadata graph. In Table~\ref{rel}, \textit{run} indicates that the user starts an execution; \textit{exe} means the execution is based on a corresponding executable file; \textit{read/write} indicates the I/O operations from executions to data objects. As all the relationships are directed, it will be difficult to travel back from \textit{dest} nodes to \textit{src} nodes. So, in current model, we define corresponding reversed relationship for each relationship to accelerate the reversed traversal (the \textit{wasXXBy} relationships).
\begin{table}[h]
\caption{Default Relationships Definition.}
@@ -35,9 +35,11 @@ There are several basic relationships between the basic entities in Table~\ref{
\end{tabular}
\end{table}
-There are several \textit{belongs/contains} relationships. In the Execution entity case, it means one job contains multiple processes, which in turns belong to this job. In the Data Objects case, it can show that one directory may contain multiple files or directories. Users can create their own relationships from two existing entities. For example, two users can have a new relationships called \textit{login-together} if they login the system roughly at the same time.
+There are several \textit{belongs/contains} relationships in Table~\ref{rel}. In the Execution entity case, it means one job contains multiple processes, which in turns belong to this job. In the Data Objects case, it captures that one directory may contain multiple files or directories.
+
+Also, users can create their own relationships based on any two existing entities. For example, two user entities can have a new relationships called \textit{login-together} if they login the system roughly at the same time.
\subsection{Property}
-Rich metadata also contain annotations on entities and their relationships. In graph model, we store them as properties, which are key-value pairs attached on vertices and edges. Users can create their own properties on existed vertices and edges except their keys need to be unique in each user's namespace. By isolating properties by users, we avoid global contention among different users. The properties could be very flexible and diverse. For example, there are properties like user name, privilege, execution parameters, file permission, and file creation and access time etc.
+Rich metadata also contain annotations on entities and their relationships. In metadata graph model, we store these as properties, which are key-value pairs attached on vertices and edges. Users can create their own properties on vertices and edges, except their keys need to be unique in each user's namespace. By isolating properties by users, we avoid global contention among different users. The properties could be very flexible and diverse. For example, there are properties like user name, privilege, execution parameters, file permission, creation time; and there are properties like data source agent, data quality score, execution environment variables, etc.
%In fact, the property can be mapped as new entity and new relationship. For example, a property of
\ No newline at end of file
diff --git a/papers/meta-graph/proto.tex b/papers/meta-graph/proto.tex
index b53405e..f02446e 100644
--- a/papers/meta-graph/proto.tex
+++ b/papers/meta-graph/proto.tex
@@ -1,17 +1,17 @@
\section{Metadata Graph Prototype}
-Collecting rich metadata usually requires modifications on HPC runtime as many rich metadata were generated from the runtime systems, like the jobs, processes, and read/write operations~\cite{muniswamy2006provenance,muniswamy2009layering, braun2006issues}. To help understand the attributes of a rich metadata graph in an HPC context, we exploit Darshan trace logs as a source of rich metadata in current prototyping~\cite{carns200924}.
+Collecting rich metadata usually requires modifications to HPC services as many rich metadata are generated from the runtime systems, like the jobs, processes, and read/write operations~\cite{muniswamy2006provenance,muniswamy2009layering, braun2006issues}. To help understand the attributes of a rich metadata graph in an HPC context, we exploit Darshan trace logs as a source of rich metadata in current prototyping~\cite{carns200924}.
\subsection{Mapping Strategy}
-Darshan utility is a MPI library that can be linked to users' applications and generates I/O behaviors logs during executing~\cite{carns2011understanding}. Each Darshan log file represents a distinct job. The log entries of a job contain the user id who started this job, the executable file that the job was based on, the parameters of this execution, some environmental variables, and most importantly, the file access statistics of each process (ranks) inside this job (MPI program). Note that, the collected Darshan traces have been anonymized, only storing the hashed values of file names, path, user names, and job names~\cite{darshanlog2013}.
+The Darshan utility is a MPI library that can be linked to users' applications and generates I/O behaviors logs during execution~\cite{carns2011understanding}. Each Darshan log file represents a distinct job. The log entries of a job contain the user id who started this job, the executable file that the job was based on, the parameters of this execution, some environment variables, and most importantly, the file access statistics of each process (ranks) inside this job (MPI program). Note that, the collected Darshan traces available on-line have been anonymized, only storing the hashed values of file names, path, user names, and job names~\cite{darshanlog2013}.
-We map Darshan logs to the metadata graph defined in Section II. Basically, each unique user id indicates a User entity, each Darshan log file represents a Job, all the ranks inside a job correspond to the Processes, and both the executables and data files are abstracted as different Data Object entities. Currently, Darshan does not capture directory structure as it only stores the hashed value of file paths, so we synthetically create the simplest directory structures: data files visited by each execution are considered under the same directory, and all these directories accessed by one user are placed under one directory for each user. Based on this mapping, we were able to import a whole year' Darshan trace (\textit{2013}) on Intrepid machine into an example graph~\cite{darshanlog2013}.
+We map Darshan logs to the metadata graph defined in Section II. Basically, each unique user id indicates a User entity, each Darshan log file represents a Job, all the ranks inside a job correspond to the Processes, and both the executables and data files are abstracted as Data Object entities. Currently, Darshan does not capture directory structure as it only stores the hashed value of file paths, so we synthetically create the simplest directory structures: data files visited by each execution are considered under the same directory, and all these directories accessed by one user are placed under one directory for each user (this structure is only possible in the graph model). Based on this mapping, we were able to import a year's Darshan traces (\textit{2013}) from the Intrepid machine into an example graph~\cite{darshanlog2013}.
%We collect basic metadata in Darshan logs as the properties too: the \textit{Type} property is set to each entity and relationship, the \textit{start\_time} and \textit{end\_time} of a job is mapped to the $start_{ts}$ and $end_{ts}$ properties of \textit{run/exe} and \textit{read/write} relationships.
-In fact, current mapping of Darshan logs emitted many common metadata due to the limitation of the data sources. For example, the logs do not have the metadata about the users; do not contain the directory structure or file permissions; and each job is based on a single executable file without configuration files and parameters. However, the generated graph still show many interesting properties of such metadata graph and offer a great potential in data management.
+%In fact, the prototype omitted many common metadata due to the limitation of the data sources. The Darshan logs do not have the metadata about the users; do not contain the directory structure or any file attributes; all jobs are based on a single executable file without parameters or configuration files. However, the generated graph still shows many interesting properties of such metadata graphs and offers good insights into the approach.
\subsection{Graph Size}
-The first property of metadata graph is the their potential size. The graph size described here is in terms of the number of vertice, edges, and properties\footnotemark. The real storage size is based on those numbers but may vary under different data structures and storage layouts.
+The first property of metadata graph is their potential size. The graph size described here is in terms of the number of vertice and edges\footnotemark. And, the real storage size is based on those numbers but may vary under different data structures and storage layouts.
\begin{table}[h]
\caption{Darshan Graph Size and Some Comparisons.}
@@ -20,23 +20,23 @@ The first property of metadata graph is the their potential size. The graph size
\begin{tabular}{|c||c|c|c|c|}
\hline
Number & Basic & \multicolumn{1}{c|}{\begin{tabular}[c]{@{}c@{}}With\\ I/O Ranks\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@{}c@{}}With\\ Full Ranks\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@{}c@{}}With\\ Directory\end{tabular}} \\ \hline
-\begin{tabular}[c]{@{}c@{}}Vertice\end{tabular} & 34,656 K & 41,729 K & 147,886 K & 147,934 K \\ \hline
-\begin{tabular}[c]{@{}c@{}}Edges\end{tabular} & 126,488 K & 133,561 K & 239,766 K & 366,253 K \\ \hline
+\begin{tabular}[c]{@{}c@{}}Vertice\end{tabular} & 34.6 M & 41.7 M & 147.8 M & 147.9 M \\ \hline
+\begin{tabular}[c]{@{}c@{}}Edges\end{tabular} & 126.5 M & 133.6 M & 239.8 M & 366.3 M \\ \hline
%\begin{tabular}[c]{@{}c@{}}Properties\end{tabular} & 448,775 K & 484,141 K & 1,015,070 K & 1,394,628 K \\ \hline
%\begin{tabular}[c]{@{}c@{}}Total Size\end{tabular} & 609,918 K & 659,431 K & 1,402,722 K & 1,908,815 K \\ \hline
\hline
\hline
- & \begin{tabular}[c]{@{}c@{}}Road Graph \\ USA~\cite{demetrescu2009shortest} \end{tabular} & \begin{tabular}[c]{@{}c@{}}Twitter\end{tabular} & \begin{tabular}[c]{@{}c@{}} Orkut~\cite{yang2012defining} \end{tabular} & \begin{tabular}[c]{@{}c@{}}Web Page \\ Graph\end{tabular} \\ \hline
-Vertice & 24 M & 645 M~\cite{twitters} & 3 M & 2.1 B~\cite{guillaume2002web} \\ \hline
-Edges & 29 M & 81,364 M~\cite{twitteredge} & 117 M & 15 B ~\cite{guillaume2002web} \\ \hline
+ & \begin{tabular}[c]{@{}c@{}}Road Graph \\ USA~\cite{demetrescu2009shortest} \end{tabular} & \begin{tabular}[c]{@{}c@{}}Twitter\end{tabular} & \begin{tabular}[c]{@{}c@{}} Facebook~\cite{facebookgs} \end{tabular} & \begin{tabular}[c]{@{}c@{}}Web Page \\ (2002)~\cite{guillaume2002web}\end{tabular} \\ \hline
+Vertice & 24 M & 645 M~\cite{twitters} & 1.28 B & 2.1 B \\ \hline
+Edges & 29 M & 81.4 B~\cite{twitteredge} & 256 B & 15 B \\ \hline
\end{tabular}
\end{table}
-\footnotetext{K = thousand; M = million; B = billion; T = trillion}
-The top half of Table~\ref{abs} shows the graph size of example metadata graph for different levels of detail. The first column considers the job as Execution entity and eliminates the all the processes (ranks) inside this job. All the I/O behaviors from different processes inside a job are considered as from the Job entity. The second column (\textit{With I/O Ranks}) records part of the ranks which have I/O operations. In many cases, this indicates the rank 0 process or the aggregators in two-phase I/O. The third column (\textit{With Full Ranks}) records all the ranks as Processes entities no matter they performed I/O or not. The last column shows the graph size with the synthetic directory structures and full ranks. This table clearly shows that increasing the level of detailed during collecting metadata will dramatically increase the graph size. So, administrators should wisely choose their metadata according to the usage. And, the user-defined relationships should be cre
ated with caution to avoid huge increasing in graph size too.
+\footnotetext{M = million; B = billion; T = trillion}
+The top half of Table~\ref{abs} shows the graph size of our example metadata graph for different levels of detail. The first column considers the job as Execution entity and eliminates all the processes (ranks) information inside this job. All the I/O behaviors from different processes inside a job are considered from the Job entity. The second column (\textit{With I/O Ranks}) records part of the ranks which have I/O operations. In many cases, this indicates the rank 0 process or the aggregators in two-phase I/O. The third column (\textit{With Full Ranks}) records all the ranks as Processes entities no matter whether they performed I/O or not. The last column shows the graph size with the complete synthetic directory structures and full ranks. This table clearly shows that increasing the level of detail will dramatically increase the graph size. %So, administrators should wisely choose their metadata according to the usage. And, the user-defined relationships should be creat
ed with caution to avoid huge increasing in graph size too.
-In the button half of Table~\ref{abs}, we show several typical large-scale graphs from different fields, including the social network (e.g. Twitter, Orkut), the road map graph (e.g. USA map), and the Internet web pages. By comparing with them, we can notice that, although the metadata graph is large and could easily be much larger, this kind of size has already been managed well in many existing systems.
+In the button half of Table~\ref{abs}, we show several typical large-scale graphs from different fields, including the social network (e.g. Twitter, Facebook), the road map graph (e.g. USA map), and the Internet web pages (estimation of 2002). By comparing with them, we note that, although the metadata graph is large, graphs of this size are already manageable in many existing systems.
%, and the largest and most complex graph, human brain%they are still manageable using current techniques.
\subsection{Graph Structure}
@@ -48,14 +48,14 @@ In addition to the graph size, the graph structure also matters in future storag
\centering
\begin{tabular}{|c|c|c|c|c|c|}
\hline
- & User & Job & Proc. & Rank & File \\ \hline
- Num & 117 & 47,592 & 10,085,931 & 113,278,038 & 34,608,033 \\ \hline
+ & Users & Jobs & Proc. & Ranks & Files \\ \hline
+ Num & 177 & 47,592 & 10,085,931 & 113,278,038 & 34,608,033 \\ \hline
\end{tabular}
\end{table}
-This table shows different entities in metadata graph have totally different size, so in this section, we will consider them separately. Figure~\ref{doe} shows the node degree distributions of those four different entities. The $x$-axis denotes the degree, the $y$-axis shows the number of vertices which have that number of degree. Both $x$-axis and $y$-axis are `log' values.
+This table shows that different entities in HPC environment may have totally different size, so the edges between them also could diverse a lot. In this section, we will consider them separately. Based on the mapping strategy, the degree of a user node indicates how many jobs each user submitted; the degree of a job node shows how many processes it contains; the degree of each process node shows how many files are read or written by it. And the file node degree shows all the data accesses from different processes including executing, reading, and writing.
-%The degree of a user node indicates how many jobs the user submitted in total; the degree of a job node shows how many processes it contains; the degree of each process node shows how many files are read or written by it. Data Objects both have \textit{contains/belongs} edges to other Data Objects and also have \textit{exe, read, write} edges to the Execution nodes.
+Figure~\ref{doe} shows the node degree distributions of four basic entities: User, Job, Process, and File (Data Object) in the metadata graphs. In these figures, all $x$-axis denotes the degree and $y$-axis shows the number of vertices which have that number of degree. Both $x$-axis and $y$-axis are $log_{10}$ values. From these figures, we can notice that all those four entities have a common attribute that most of entities have very small degrees and a small number of entities have much larger degrees. Take \textit{User} node as an example, there are totally 177 users who submitted jobs in this year (\textit{2013}). Among them, around 10\% users actually submitted more than 80\% of all the jobs. For other three entities, we can observe the similar phenomenon.
\begin{Figure}{Node degree for different entities.}[doe]
\graphfile*[3]{exps/udegree.pdf}[User node degree]
@@ -63,17 +63,19 @@ This table shows different entities in metadata graph have totally different siz
\graphfile*[3]{exps/pdegree.pdf}[Process node degree]
\graphfile*[3]{exps/fdegree.pdf}[File node degree]
\end{Figure}
-
+
-%To describe the graph structure, we compared them with
-In nature, many graphs obey the \textit{skewed} power-law degree distribution~\cite{faloutsos1999power}. It means most vertices have relatively few neighbors while a few vertices have many neighbors. We can use function $\textbf{P}(d) \propto d^{-\alpha}$ to describe the probability that a vertex has degree $d$ in such graphs. Here, $\alpha$ is a positive constant that control the ``skewness'' of the degree distribution: higher $\alpha$ indicates lower density, which means vast majority of vertices are low degree. Lower $\alpha$ shows higher density and more high degree vertices.
+Actually, in nature, many graphs have this similar attribute. They can be described using the \textit{skewed} power-law degree distribution~\cite{faloutsos1999power}, which means, in these graphs, most vertices have relatively few neighbors while a few vertices have much more neighbors. We use function $\textbf{P}(d) \propto d^{-\alpha}$ to describe the probability that a vertex has degree $d$ in such graphs. Here, $\alpha$ is a constant parameter of the distribution known as the \textit{exponent} or \textit{scaling parameter}.
+%It typically lies in the range 2 $< \alpha<$ 3.
-To compare the actual distribution with the power-law attributes, we plot three lines of the power-law distribution with different $\alpha$ values. Intuitively, we can notice that the \textit{File} nodes fit the power-law degree distribution best as Fig.~\ref{doe}(d) shows. The degree of a file node indicates the number of processes that ever read/write it. So, the distribution indicates that most files are seldom accessed, and a very small part of files are visited highly frequently. Second, as the larger $\alpha$ value ($2 \to 4$) actually fits the distribution better, which indicates the \textit{File} nodes have lower density, saying that the majority of files have low degrees.
+According to Fig.~\ref{doe}, we speculate some of the entities fit the power-law distribution. So, we plot three lines of the power-law distribution with different $\alpha$ values in each figure (we eliminate \textit{User} as there are only 177 user samples, which is too small to tell the distribution.). Intuitively, we can notice that the \textit{File} nodes come closer to the power-law degree distribution as Fig.~\ref{doe}(d) shows. This is reasonable as this data locality attribute, which indicates that most files are seldom accessed, and a very small part of files are visited highly frequently, also was found in many other systems. At the same time, other two distributions, including \textit{Process} and \textit{Job}, are not visually obvious to fit the power-law distribution. Like, for the \textit{Process} (Fig.~\ref{doe}(c)) , there are not enough low degree processes nodes, and, for \textit{Job} (Fig.~\ref{doe}(b)), the distribution scatters more randomly than other t
wo.
+%which indicates most processes in HPC systems tend to visit at least several files
+%between the $\alpha=3$ line and $\alpha=0.5$ line.
+%This is reasonable since most jobs in HPC tend to have some number of processes for parallelism. randomly
-%This is reasonable since most jobs in HPC tend to have more processes
-Other three distributions of \textit{Process}, \textit{Job} and \textit{User}, are more and more unlike the power-law distribution. For the \textit{Process} graph (Fig.~\ref{doe}(c)) , there are not enough processes with low degree. This means most processes in HPC systems tend to visit at least several files. For \textit{Job} graph (Fig.~\ref{doe}(b)), the distribution scatters randomly between the $\alpha=3$ line and $\alpha=0.5$ line. As there are only 117 users, we do not consider they fit any distribution. But from Fig~\ref{doe}(a), we still can see most users only issue several jobs, a very small number of users will issue the most jobs.
+% But from Fig~\ref{doe}(a), we still can see most users only issue several jobs, a very small number of users will issue the most jobs.
-These graph structures direct the way of graph storage and processing including graph partition and storage layout. Moreover, they also lay as a basis to create synthetic graphs for evaluation test as collecting enough rich metadata in running HPC systems are still hard.
+We are still checking whether all metadata graphs generated from different use cases will still fit the power-law distribution and finding out whether adding/deleting nodes/edges will change this attribute based on the statistical analytical strategy introduced by A. Clauset et. al.~\cite{clauset2009power}. The preliminary results still show a strong indication that the power-law distribution is an accurate estimation for these basic entities in metadata graphs. This result plays an important role in our future designing in graph partition and storage strategy. Moreover, it also serves as a basis for creating synthetic graphs for evaluation since collecting large-scale rich metadata in running HPC systems is still challenging. %We are still exploring the This is still a undergoing work due to the complexity of power-law distribution and the flexibility of metadata graphs.
%To calculate the graph diameter and connected component, different entities and relationships are considered as the same kind of nodes and edges. For our example graph, the diameter is $x$ and the connected component number is $y$. This value can vary while we define new entity or relationship as next subsection shows.
diff --git a/papers/meta-graph/exps/fdegree.pdf b/papers/meta-graph/report/exps/fdegree.pdf
similarity index 95%
copy from papers/meta-graph/exps/fdegree.pdf
copy to papers/meta-graph/report/exps/fdegree.pdf
index 4ca0abf..2911b4b 100644
Binary files a/papers/meta-graph/exps/fdegree.pdf and b/papers/meta-graph/report/exps/fdegree.pdf differ
diff --git a/papers/meta-graph/exps/udegree.pdf b/papers/meta-graph/report/exps/histplot.pdf
similarity index 52%
copy from papers/meta-graph/exps/udegree.pdf
copy to papers/meta-graph/report/exps/histplot.pdf
index 127a3c0..d01297e 100644
Binary files a/papers/meta-graph/exps/udegree.pdf and b/papers/meta-graph/report/exps/histplot.pdf differ
diff --git a/papers/meta-graph/exps/jdegree.pdf b/papers/meta-graph/report/exps/jdegree.pdf
similarity index 80%
copy from papers/meta-graph/exps/jdegree.pdf
copy to papers/meta-graph/report/exps/jdegree.pdf
index 2e379e7..2408949 100644
Binary files a/papers/meta-graph/exps/jdegree.pdf and b/papers/meta-graph/report/exps/jdegree.pdf differ
diff --git a/papers/meta-graph/exps/pdegree.pdf b/papers/meta-graph/report/exps/pdegree.pdf
similarity index 97%
copy from papers/meta-graph/exps/pdegree.pdf
copy to papers/meta-graph/report/exps/pdegree.pdf
index a2b5aec..06457b5 100644
Binary files a/papers/meta-graph/exps/pdegree.pdf and b/papers/meta-graph/report/exps/pdegree.pdf differ
diff --git a/papers/meta-graph/report/exps/poweRlaw-cdf.pdf b/papers/meta-graph/report/exps/poweRlaw-cdf.pdf
new file mode 100644
index 0000000..4f5bc1a
Binary files /dev/null and b/papers/meta-graph/report/exps/poweRlaw-cdf.pdf differ
diff --git a/papers/meta-graph/exps/udegree.pdf b/papers/meta-graph/report/exps/udegree.pdf
similarity index 75%
copy from papers/meta-graph/exps/udegree.pdf
copy to papers/meta-graph/report/exps/udegree.pdf
index 127a3c0..b7515c7 100644
Binary files a/papers/meta-graph/exps/udegree.pdf and b/papers/meta-graph/report/exps/udegree.pdf differ
diff --git a/papers/meta-graph/exps/udegree.pdf b/papers/meta-graph/report/exps/user-histogram.pdf
similarity index 57%
copy from papers/meta-graph/exps/udegree.pdf
copy to papers/meta-graph/report/exps/user-histogram.pdf
index 127a3c0..2b375cc 100644
Binary files a/papers/meta-graph/exps/udegree.pdf and b/papers/meta-graph/report/exps/user-histogram.pdf differ
diff --git a/papers/meta-graph/report/power-law-dist.aux b/papers/meta-graph/report/power-law-dist.aux
new file mode 100644
index 0000000..f8dd830
--- /dev/null
+++ b/papers/meta-graph/report/power-law-dist.aux
@@ -0,0 +1,32 @@
+\relax
+\@writefile{toc}{\contentsline {section}{\numberline {1}Background}{1}}
+\@writefile{lot}{\contentsline {table}{\numberline {1}{\ignorespaces Statistics of Metadata Graph.}}{1}}
+\newlabel{t2}{{1}{1}}
+\@writefile{toc}{\contentsline {section}{\numberline {2}Fist Step}{1}}
+\newlabel{doe:a}{{1a}{1}}
+\newlabel{sub@doe:a}{{a}{1}}
+\newlabel{doe:b}{{1b}{1}}
+\newlabel{sub@doe:b}{{b}{1}}
+\newlabel{doe:c}{{1c}{1}}
+\newlabel{sub@doe:c}{{c}{1}}
+\newlabel{doe:d}{{1d}{1}}
+\newlabel{sub@doe:d}{{d}{1}}
+\@writefile{lof}{\contentsline {figure}{\numberline {1}{\ignorespaces Node degree for different entities.}}{1}}
+\@writefile{lof}{\contentsline {subfigure}{\numberline{a}{\ignorespaces {User node degree}}}{1}}
+\@writefile{lof}{\contentsline {subfigure}{\numberline{b}{\ignorespaces {Job node degree}}}{1}}
+\@writefile{lof}{\contentsline {subfigure}{\numberline{c}{\ignorespaces {Process node degree}}}{1}}
+\@writefile{lof}{\contentsline {subfigure}{\numberline{d}{\ignorespaces {File node degree}}}{1}}
+\newlabel{doe}{{1}{1}}
+\@writefile{toc}{\contentsline {section}{\numberline {3}Some Basis on Power-Law Distribution}{1}}
+\newlabel{cdf}{{4}{2}}
+\@writefile{toc}{\contentsline {section}{\numberline {4}Histogram Check}{2}}
+\@writefile{lof}{\contentsline {figure}{\numberline {2}{\ignorespaces Histogram of User, Process, File Degree, and an example power-law dist.}}{2}}
+\newlabel{f1}{{2}{2}}
+\@writefile{lof}{\contentsline {figure}{\numberline {3}{\ignorespaces Histogram of User with different intervals.}}{2}}
+\newlabel{f2}{{3}{2}}
+\@writefile{lot}{\contentsline {table}{\numberline {2}{\ignorespaces Fitted parameters for different entities and different distributions.}}{3}}
+\newlabel{t1}{{2}{3}}
+\@writefile{toc}{\contentsline {section}{\numberline {5}Using PoweRlaw Package}{3}}
+\@writefile{lof}{\contentsline {figure}{\numberline {4}{\ignorespaces Degree CDF for different entities. The \textbf {red line} denotes Power-Law distribution, the \textbf {blue line} shows the Log-Normal distribution.}}{3}}
+\newlabel{f3}{{4}{3}}
+\@writefile{toc}{\contentsline {section}{\numberline {6}Conclusion}{4}}
diff --git a/papers/meta-graph/report/power-law-dist.log b/papers/meta-graph/report/power-law-dist.log
new file mode 100644
index 0000000..63ae5c7
--- /dev/null
+++ b/papers/meta-graph/report/power-law-dist.log
@@ -0,0 +1,337 @@
+This is pdfTeX, Version 3.1415926-2.5-1.40.14 (TeX Live 2013) (format=pdflatex 2013.5.30) 27 AUG 2014 15:22
+entering extended mode
+ restricted \write18 enabled.
+ %&-line parsing enabled.
+**power-law-dist.tex
+(./power-law-dist.tex
+LaTeX2e <2011/06/27>
+Babel <3.9f> and hyphenation patterns for 78 languages loaded.
+(/usr/local/texlive/2013/texmf-dist/tex/latex/base/article.cls
+Document Class: article 2007/10/19 v1.4h Standard LaTeX document class
+(/usr/local/texlive/2013/texmf-dist/tex/latex/base/size10.clo
+File: size10.clo 2007/10/19 v1.4h Standard LaTeX file (size option)
+)
+\c@part=\count79
+\c@section=\count80
+\c@subsection=\count81
+\c@subsubsection=\count82
+\c@paragraph=\count83
+\c@subparagraph=\count84
+\c@figure=\count85
+\c@table=\count86
+\abovecaptionskip=\skip41
+\belowcaptionskip=\skip42
+\bibindent=\dimen102
+) (./usenix.sty
+(/usr/local/texlive/2013/texmf-dist/tex/latex/psnfss/mathptmx.sty
+Package: mathptmx 2005/04/12 PSNFSS-v9.2a Times w/ Math, improved (SPQR, WaS)
+LaTeX Font Info: Redeclaring symbol font `operators' on input line 28.
+LaTeX Font Info: Overwriting symbol font `operators' in version `normal'
+(Font) OT1/cmr/m/n --> OT1/ztmcm/m/n on input line 28.
+LaTeX Font Info: Overwriting symbol font `operators' in version `bold'
+(Font) OT1/cmr/bx/n --> OT1/ztmcm/m/n on input line 28.
+LaTeX Font Info: Redeclaring symbol font `letters' on input line 29.
+LaTeX Font Info: Overwriting symbol font `letters' in version `normal'
+(Font) OML/cmm/m/it --> OML/ztmcm/m/it on input line 29.
+LaTeX Font Info: Overwriting symbol font `letters' in version `bold'
+(Font) OML/cmm/b/it --> OML/ztmcm/m/it on input line 29.
+LaTeX Font Info: Redeclaring symbol font `symbols' on input line 30.
+LaTeX Font Info: Overwriting symbol font `symbols' in version `normal'
+(Font) OMS/cmsy/m/n --> OMS/ztmcm/m/n on input line 30.
+LaTeX Font Info: Overwriting symbol font `symbols' in version `bold'
+(Font) OMS/cmsy/b/n --> OMS/ztmcm/m/n on input line 30.
+LaTeX Font Info: Redeclaring symbol font `largesymbols' on input line 31.
+LaTeX Font Info: Overwriting symbol font `largesymbols' in version `normal'
+(Font) OMX/cmex/m/n --> OMX/ztmcm/m/n on input line 31.
+LaTeX Font Info: Overwriting symbol font `largesymbols' in version `bold'
+(Font) OMX/cmex/m/n --> OMX/ztmcm/m/n on input line 31.
+\symbold=\mathgroup4
+\symitalic=\mathgroup5
+LaTeX Font Info: Redeclaring math alphabet \mathbf on input line 34.
+LaTeX Font Info: Overwriting math alphabet `\mathbf' in version `normal'
+(Font) OT1/cmr/bx/n --> OT1/ptm/bx/n on input line 34.
+LaTeX Font Info: Overwriting math alphabet `\mathbf' in version `bold'
+(Font) OT1/cmr/bx/n --> OT1/ptm/bx/n on input line 34.
+LaTeX Font Info: Redeclaring math alphabet \mathit on input line 35.
+LaTeX Font Info: Overwriting math alphabet `\mathit' in version `normal'
+(Font) OT1/cmr/m/it --> OT1/ptm/m/it on input line 35.
+LaTeX Font Info: Overwriting math alphabet `\mathit' in version `bold'
+(Font) OT1/cmr/bx/it --> OT1/ptm/m/it on input line 35.
+LaTeX Info: Redefining \hbar on input line 50.
+))
+(/usr/local/texlive/2013/texmf-dist/tex/latex/graphics/epsfig.sty
+Package: epsfig 1999/02/16 v1.7a (e)psfig emulation (SPQR)
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/graphics/graphicx.sty
+Package: graphicx 1999/02/16 v1.0f Enhanced LaTeX Graphics (DPC,SPQR)
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/graphics/keyval.sty
+Package: keyval 1999/03/16 v1.13 key=value parser (DPC)
+\KV@toks@=\toks14
+)
+(/usr/local/texlive/2013/texmf-dist/tex/latex/graphics/graphics.sty
+Package: graphics 2009/02/05 v1.0o Standard LaTeX Graphics (DPC,SPQR)
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/graphics/trig.sty
+Package: trig 1999/03/16 v1.09 sin cos tan (DPC)
+)
+(/usr/local/texlive/2013/texmf-dist/tex/latex/latexconfig/graphics.cfg
+File: graphics.cfg 2010/04/23 v1.9 graphics configuration of TeX Live
+)
+Package graphics Info: Driver file: pdftex.def on input line 91.
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/pdftex-def/pdftex.def
+File: pdftex.def 2011/05/27 v0.06d Graphics/color for pdfTeX
+
+(/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/infwarerr.sty
+Package: infwarerr 2010/04/08 v1.3 Providing info/warning/error messages (HO)
+)
+(/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/ltxcmds.sty
+Package: ltxcmds 2011/11/09 v1.22 LaTeX kernel commands for general use (HO)
+)
+\Gread@gobject=\count87
+))
+\Gin@req@height=\dimen103
+\Gin@req@width=\dimen104
+)
+\epsfxsize=\dimen105
+\epsfysize=\dimen106
+)
+(/usr/local/texlive/2013/texmf-dist/tex/latex/endnotes/endnotes.sty
+\c@endnote=\count88
+\endnotesep=\dimen107
+\@enotes=\write3
+)
+(/usr/local/texlive/2013/texmf-dist/tex/latex/bosisio/graphfig.sty
+Package: graphfig 1997/12/15 v2.2 Commands to include graphics files (FB)
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/subfigure/subfigure.sty
+Package: subfigure 2002/03/15 v2.1.5 subfigure package
+\subfigtopskip=\skip43
+\subfigcapskip=\skip44
+\subfigcaptopadj=\dimen108
+\subfigbottomskip=\skip45
+\subfigcapmargin=\dimen109
+\subfiglabelskip=\skip46
+\c@subfigure=\count89
+\c@lofdepth=\count90
+\c@subtable=\count91
+\c@lotdepth=\count92
+
+****************************************
+* Local config file subfigure.cfg used *
+****************************************
+(/usr/local/texlive/2013/texmf-dist/tex/latex/subfigure/subfigure.cfg)
+\subfig@top=\skip47
+\subfig@bottom=\skip48
+))
+(./power-law-dist.aux)
+\openout1 = `power-law-dist.aux'.
+
+LaTeX Font Info: Checking defaults for OML/cmm/m/it on input line 9.
+LaTeX Font Info: ... okay on input line 9.
+LaTeX Font Info: Checking defaults for T1/cmr/m/n on input line 9.
+LaTeX Font Info: ... okay on input line 9.
+LaTeX Font Info: Checking defaults for OT1/cmr/m/n on input line 9.
+LaTeX Font Info: ... okay on input line 9.
+LaTeX Font Info: Checking defaults for OMS/cmsy/m/n on input line 9.
+LaTeX Font Info: ... okay on input line 9.
+LaTeX Font Info: Checking defaults for OMX/cmex/m/n on input line 9.
+LaTeX Font Info: ... okay on input line 9.
+LaTeX Font Info: Checking defaults for U/cmr/m/n on input line 9.
+LaTeX Font Info: ... okay on input line 9.
+LaTeX Font Info: Try loading font information for OT1+ptm on input line 9.
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/psnfss/ot1ptm.fd
+File: ot1ptm.fd 2001/06/04 font definitions for OT1/ptm.
+)
+\big@size=\dimen110
+
+(/usr/local/texlive/2013/texmf-dist/tex/context/base/supp-pdf.mkii
+[Loading MPS to PDF converter (version 2006.09.02).]
+\scratchcounter=\count93
+\scratchdimen=\dimen111
+\scratchbox=\box26
+\nofMPsegments=\count94
+\nofMParguments=\count95
+\everyMPshowfont=\toks15
+\MPscratchCnt=\count96
+\MPscratchDim=\dimen112
+\MPnumerator=\count97
+\makeMPintoPDFobject=\count98
+\everyMPtoPDFconversion=\toks16
+) (/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/pdftexcmds.sty
+Package: pdftexcmds 2011/11/29 v0.20 Utility functions of pdfTeX for LuaTeX (HO
+)
+
+(/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/ifluatex.sty
+Package: ifluatex 2010/03/01 v1.3 Provides the ifluatex switch (HO)
+Package ifluatex Info: LuaTeX not detected.
+)
+(/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/ifpdf.sty
+Package: ifpdf 2011/01/30 v2.3 Provides the ifpdf switch (HO)
+Package ifpdf Info: pdfTeX in PDF mode is detected.
+)
+Package pdftexcmds Info: LuaTeX not detected.
+Package pdftexcmds Info: \pdf@primitive is available.
+Package pdftexcmds Info: \pdf@ifprimitive is available.
+Package pdftexcmds Info: \pdfdraftmode found.
+)
+(/usr/local/texlive/2013/texmf-dist/tex/latex/oberdiek/epstopdf-base.sty
+Package: epstopdf-base 2010/02/09 v2.5 Base part for package epstopdf
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/oberdiek/grfext.sty
+Package: grfext 2010/08/19 v1.1 Manage graphics extensions (HO)
+
+(/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/kvdefinekeys.sty
+Package: kvdefinekeys 2011/04/07 v1.3 Define keys (HO)
+))
+(/usr/local/texlive/2013/texmf-dist/tex/latex/oberdiek/kvoptions.sty
+Package: kvoptions 2011/06/30 v3.11 Key value format for package options (HO)
+
+(/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/kvsetkeys.sty
+Package: kvsetkeys 2012/04/25 v1.16 Key value parser (HO)
+
+(/usr/local/texlive/2013/texmf-dist/tex/generic/oberdiek/etexcmds.sty
+Package: etexcmds 2011/02/16 v1.5 Avoid name clashes with e-TeX commands (HO)
+Package etexcmds Info: Could not find \expanded.
+(etexcmds) That can mean that you are not using pdfTeX 1.50 or
+(etexcmds) that some package has redefined \expanded.
+(etexcmds) In the latter case, load this package earlier.
+)))
+Package grfext Info: Graphics extension search list:
+(grfext) [.png,.pdf,.jpg,.mps,.jpeg,.jbig2,.jb2,.PNG,.PDF,.JPG,.JPE
+G,.JBIG2,.JB2,.eps]
+(grfext) \AppendGraphicsExtensions on input line 452.
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/latexconfig/epstopdf-sys.cfg
+File: epstopdf-sys.cfg 2010/07/13 v1.3 Configuration of (r)epstopdf for TeX Liv
+e
+))
+LaTeX Font Info: Font shape `OT1/ptm/bx/n' in size <14.4> not available
+(Font) Font shape `OT1/ptm/b/n' tried instead on input line 13.
+LaTeX Font Info: Try loading font information for OT1+ztmcm on input line 13
+.
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/psnfss/ot1ztmcm.fd
+File: ot1ztmcm.fd 2000/01/03 Fontinst v1.801 font definitions for OT1/ztmcm.
+)
+LaTeX Font Info: Try loading font information for OML+ztmcm on input line 13
+.
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/psnfss/omlztmcm.fd
+File: omlztmcm.fd 2000/01/03 Fontinst v1.801 font definitions for OML/ztmcm.
+)
+LaTeX Font Info: Try loading font information for OMS+ztmcm on input line 13
+.
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/psnfss/omsztmcm.fd
+File: omsztmcm.fd 2000/01/03 Fontinst v1.801 font definitions for OMS/ztmcm.
+)
+LaTeX Font Info: Try loading font information for OMX+ztmcm on input line 13
+.
+
+(/usr/local/texlive/2013/texmf-dist/tex/latex/psnfss/omxztmcm.fd
+File: omxztmcm.fd 2000/01/03 Fontinst v1.801 font definitions for OMX/ztmcm.
+)
+LaTeX Font Info: Font shape `OT1/ptm/bx/n' in size <12> not available
+(Font) Font shape `OT1/ptm/b/n' tried instead on input line 13.
+LaTeX Font Info: Font shape `OT1/ptm/bx/n' in size <9> not available
+(Font) Font shape `OT1/ptm/b/n' tried instead on input line 13.
+LaTeX Font Info: Font shape `OT1/ptm/bx/n' in size <7> not available
+(Font) Font shape `OT1/ptm/b/n' tried instead on input line 13.
+LaTeX Font Info: Font shape `OT1/ptm/bx/n' in size <10> not available
+(Font) Font shape `OT1/ptm/b/n' tried instead on input line 21.
+LaTeX Font Info: Font shape `OT1/ptm/bx/n' in size <7.4> not available
+(Font) Font shape `OT1/ptm/b/n' tried instead on input line 21.
+LaTeX Font Info: Font shape `OT1/ptm/bx/n' in size <6> not available
+(Font) Font shape `OT1/ptm/b/n' tried instead on input line 21.
+
+<exps/udegree.pdf, id=1, 578.16pt x 361.35pt>
+File: exps/udegree.pdf Graphic file (type pdf)
+ <use exps/udegree.pdf>
+Package pdftex.def Info: exps/udegree.pdf used on input line 35.
+(pdftex.def) Requested size: 578.15858pt x 361.34912pt.
+
+<exps/jdegree.pdf, id=3, 578.16pt x 361.35pt>
+File: exps/jdegree.pdf Graphic file (type pdf)
+ <use exps/jdegree.pdf>
+Package pdftex.def Info: exps/jdegree.pdf used on input line 35.
+(pdftex.def) Requested size: 578.15858pt x 361.34912pt.
+
+<exps/pdegree.pdf, id=5, 578.16pt x 361.35pt>
+File: exps/pdegree.pdf Graphic file (type pdf)
+ <use exps/pdegree.pdf>
+Package pdftex.def Info: exps/pdegree.pdf used on input line 36.
+(pdftex.def) Requested size: 578.15858pt x 361.34912pt.
+
+<exps/fdegree.pdf, id=7, 578.16pt x 361.35pt>
+File: exps/fdegree.pdf Graphic file (type pdf)
+ <use exps/fdegree.pdf>
+Package pdftex.def Info: exps/fdegree.pdf used on input line 37.
+(pdftex.def) Requested size: 578.15858pt x 361.34912pt.
+ [1{/usr/local/texlive/2013/texmf-var/fonts/map/pdftex/updmap/pdftex.map}
+
+
+ <./exps/udegree.pdf> <./exps/jdegree.pdf> <./exps/pdegree.pdf> <./exps/fdegree
+.pdf>]
+<exps/histplot.pdf, id=52, 483.8075pt x 347.2975pt>
+File: exps/histplot.pdf Graphic file (type pdf)
+ <use exps/histplot.pdf>
+Package pdftex.def Info: exps/histplot.pdf used on input line 85.
+(pdftex.def) Requested size: 231.26378pt x 166.012pt.
+
+Overfull \hbox (5.42003pt too wide) in paragraph at lines 85--86
+ []
+ []
+
+LaTeX Font Info: Font shape `OT1/ptm/bx/it' in size <10> not available
+(Font) Font shape `OT1/ptm/b/it' tried instead on input line 99.
+<exps/user-histogram.pdf, id=53, 483.8075pt x 347.2975pt>
+File: exps/user-histogram.pdf Graphic file (type pdf)
+
+<use exps/user-histogram.pdf>
+Package pdftex.def Info: exps/user-histogram.pdf used on input line 111.
+(pdftex.def) Requested size: 216.81pt x 155.63591pt.
+ [2 <./exps/histplot.pdf> <./exps/user-histogram.pdf>] <exps/poweRlaw-cdf.pdf,
+id=77, 483.8075pt x 347.2975pt>
+File: exps/poweRlaw-cdf.pdf Graphic file (type pdf)
+
+<use exps/poweRlaw-cdf.pdf>
+Package pdftex.def Info: exps/poweRlaw-cdf.pdf used on input line 129.
+(pdftex.def) Requested size: 231.26378pt x 166.012pt.
+
+Overfull \hbox (5.42003pt too wide) in paragraph at lines 129--130
+ []
+ []
+
+
+LaTeX Warning: `!h' float specifier changed to `!ht'.
+
+[3 <./exps/poweRlaw-cdf.pdf>] [4
+
+] (./power-law-dist.aux) )
+Here is how much of TeX's memory you used:
+ 1968 strings out of 493315
+ 28763 string characters out of 6137904
+ 91466 words of memory out of 5000000
+ 5367 multiletter control sequences out of 15000+600000
+ 36141 words of font info for 76 fonts, out of 8000000 for 9000
+ 957 hyphenation exceptions out of 8191
+ 38i,13n,30p,808b,303s stack positions out of 5000i,500n,10000p,200000b,80000s
+{/usr/local/texlive/2013/texmf-dist/fonts/enc/dvips/b
+ase/8r.enc}</usr/local/texlive/2013/texmf-dist/fonts/type1/public/amsfonts/cm/c
+mmi10.pfb></usr/local/texlive/2013/texmf-dist/fonts/type1/public/amsfonts/cm/cm
+r10.pfb></usr/local/texlive/2013/texmf-dist/fonts/type1/public/amsfonts/cm/cmsy
+10.pfb></usr/local/texlive/2013/texmf-dist/fonts/type1/urw/symbol/usyr.pfb></us
+r/local/texlive/2013/texmf-dist/fonts/type1/urw/symbol/usyr.pfb></usr/local/tex
+live/2013/texmf-dist/fonts/type1/urw/times/utmb8a.pfb></usr/local/texlive/2013/
+texmf-dist/fonts/type1/urw/times/utmr8a.pfb></usr/local/texlive/2013/texmf-dist
+/fonts/type1/urw/times/utmri8a.pfb>
+Output written on power-law-dist.pdf (4 pages, 433924 bytes).
+PDF statistics:
+ 117 PDF objects out of 1000 (max. 8388607)
+ 84 compressed objects within 1 object stream
+ 0 named destinations out of 1000 (max. 500000)
+ 60 words of extra memory for PDF output out of 10000 (max. 10000000)
+
diff --git a/papers/meta-graph/report/power-law-dist.pdf b/papers/meta-graph/report/power-law-dist.pdf
new file mode 100644
index 0000000..22eaf05
Binary files /dev/null and b/papers/meta-graph/report/power-law-dist.pdf differ
diff --git a/papers/meta-graph/report/power-law-dist.synctex.gz b/papers/meta-graph/report/power-law-dist.synctex.gz
new file mode 100644
index 0000000..7774431
Binary files /dev/null and b/papers/meta-graph/report/power-law-dist.synctex.gz differ
diff --git a/papers/meta-graph/report/power-law-dist.tex b/papers/meta-graph/report/power-law-dist.tex
new file mode 100644
index 0000000..7ae38b9
--- /dev/null
+++ b/papers/meta-graph/report/power-law-dist.tex
@@ -0,0 +1,155 @@
+%\documentclass[10pt,a4paper]{report}
+\documentclass[letterpaper,twocolumn,10pt]{article}
+\usepackage{usenix,epsfig,endnotes}
+\usepackage[subfigure]{graphfig}
+\usepackage{graphicx}
+%\usepackage{draftwatermark}
+%\SetWatermarkScale{4}
+
+\begin{document}
+\title{\Large \bf Discussion on Metadata Graph Entity Degree Distribution}
+\date{8/27/2014}
+\author{Dong Dai}
+\maketitle
+\section{Background}
+As we have introduced, based on the definition of metadata graph and the Darshan trace logs, we can import one year's Darshan trace into a graph. Table~\ref{t2} shows some basic statistics of this Darshan metadata graph:
+
+\begin{table}[h]
+\caption{Statistics of Metadata Graph.}
+ \label{t2}
+\centering
+\begin{tabular}{|c|c|c|c|c|}
+\hline
+ Users & Jobs & Proc. & Ranks & Files \\ \hline
+ 177 & 47 K & 10,085 K & 113,278 K & 34,608 K \\ \hline
+\end{tabular}
+\end{table}
+
+It is not hard to tell that different entities (users, jobs, and files etc.) have totally different size, so the edges connected between them should also have different numbers. This will significantly affect the degrees of different nodes (in this document, \textit{node degree} contains both the in-edge degree and out-edge degree). So, we will check them separately.
+
+\section{Fist Step}
+The very first step would be plotting the node degree distribution for different entities, and observe what do they look like. In Fig.~\ref{doe}, we have four doubly logarithm figures for each of defined entities. All the $x$-axis denotes the degree, the $y$-axis shows the number of vertices which have that number of degree. From these figures, we can notice that all those four entities have a common attribute that most of entities have very small degrees and a small number of entities have much larger degrees.
+
+\begin{Figure}{Node degree for different entities.}[doe]
+ \graphfile*[3]{exps/udegree.pdf}[User node degree]
+ \graphfile*[3]{exps/jdegree.pdf}[Job node degree]\\
+ \graphfile*[3]{exps/pdegree.pdf}[Process node degree]
+ \graphfile*[3]{exps/fdegree.pdf}[File node degree]
+ \end{Figure}
+
+Actually, in nature, many graphs have this attribute, and they were described as following the \textit{skewed} power-law degree distributions, which indicate most vertices have relatively few neighbors while a few vertices have many neighbors. We use function $\textbf{P}(d) \propto d^{-\alpha}$ to describe the probability that a vertex has degree $d$ in such graphs. Here, $\alpha$ is a constant parameter of the distribution known as the \textit{exponent} or \textit{scaling parameter}.
+
+We speculate these entities actually fit the power-law distribution. So, we plot three lines representing the power-law distribution with different $\alpha$ values in each figure. Intuitively, we can notice that the \textit{File} nodes come closer to the power-law degree distribution as Fig.~\ref{doe}(d) shows. Other distributions including \textit{Process} and \textit{Job} are more and more unlike the power-law distribution comparing with \textit{File} nodes.
+
+The question is \textit{whether they actually fit the power-law distribution or not?} This is what we gonna try to discuss in this document.
+
+\section{Some Basis on Power-Law Distribution}
+
+Let $x$ represent the quantity, a discrete power-law distribution is described by a probability density $p(x)$ such that
+
+\begin{equation}
+p(x) = Pr(X=x)=Cx^{-\alpha}
+\end{equation}
+
+where C is a constant and $\alpha$ is the constant parameter of the distribution known as \textit{scaling parameter}. This distribution diverges at $x \to 0$, so there must be a lower bound $x_{min} > 0$ on the power-law behavior. With this $x_{min}$, we can calculate the constant factor C and turn the density function like this:
+\begin{equation}
+p(x)=\frac{x^{-\alpha}}{\zeta(\alpha, x_{min})}
+\end{equation}
+where
+\begin{equation}
+\zeta(\alpha, x_{min}) = \sum_{n=0}^{\infty} (n+x_{min})^{-\alpha}
+\end{equation}
+
+So, to describe a practical power-law distribution, there are two parameters needed: $x_{min}$ denotes the start $x$ and $\alpha$ shows the scaling. One step further, in many cases, it is useful to consider the \textit{complementary cumulative distribution function} or CDF of a power-law variable, which we denote as $P(x)$ defined to be $P(x)=Pr(X \geq x)$. In the discrete case, it will be:
+\begin{equation}
+\label{cdf}
+P(x)=\frac{\zeta(\alpha, x)}{\zeta(\alpha, x_{min})}
+\end{equation}
+
+Moreover, for many distributions we usually want to know the quantile value $\xi_{q}$. Here, $\xi_{q}$ indicates the $q_{th}$ quantile of a batch of $n$ numbers such that a fraction $q * n$ of the sample is less than $\xi_{q}$. The best know quantile is the median, $\xi_{0.5}$, which is located in the middle of the sample. Mathematically, $\xi_{q}$ actually denotes the inverse of CDF function. As we have defined CDF function in Equation~\ref{cdf}, to get the $\xi_{q}$, we need to solve this equation:
+\begin{equation}
+P(\xi_{q})=q
+\end{equation}
+
+That is, the probability of a sample is less than $\xi_{q}$ is in fact just $q$. In fact, it is very difficult to calculate $\xi_{q}$ for power-law distribution as we need to get the inverse function of $P(x)$. Based on Equation~\ref{cdf} shows, the computation would be impossible considering the wide value range of the two parameters $x_{min}$ and $\alpha$ in power-law distributions.
+
+This explains why most of the existed power-law analysis tool chains do not provide such functionalities. And also this is the reason that we do not deploy the \textit{Q-Q plot} strategy to compare distributions in this document.
+
+To fit power-law distribution to empirical data sample is not an easy job, partly due to the complexity of power-law distribution itself, and partly due to the interference from other similar distributions, like \textit{Log-Normal} or \textit{Exponent} distributions. In this document, we use two steps to show how to map our data samples to power-law distributions and discuss the plausibility of our bold assertion, which says that all four entities fit power-law distribution.
+
+\section{Histogram Check}
+The very simple way to probe for power-law behavior is to measure the quantity of interest $x$, construct a histogram representing its frequency distribution, and plot that histogram on doubly logarithmic axes. If in so doing one discovers a distribution that approximately falls on a straight line, then one can, if one is feeling particularly bold, assert that the distribution follows a power law.
+
+\begin{figure}[h!]
+\begin{center}
+ \includegraphics[width=3.2in]{exps/histplot.pdf}
+ \caption{Histogram of User, Process, File Degree, and an example power-law dist.}
+ \label{f1}
+ \end{center}
+\end{figure}
+
+From Fig~\ref{f1}, it is not too bold to say that the degrees of file nodes fit the power-law distribution since there is a clear straight line visually. However, for users and processes, their histograms are not similar with the standard histogram. So, we can not assert they ideally fit the power-law. But, we also can not assert they do not fit the power-law for two reasons.
+
+\begin{table*}
+\caption{Fitted parameters for different entities and different distributions.}
+ \label{t1}
+\centering
+\begin{tabular}{|c|c|c|c|c|c|c|}
+\hline
+\textit{\textbf{}} & \multicolumn{2}{c|}{\textbf{\begin{tabular}[c]{@{}c@{}}Power-Law\\ Params.\end{tabular}}} & \multicolumn{3}{c|}{\textbf{\begin{tabular}[c]{@{}c@{}}Log-Normal\\ Params.\end{tabular}}} & \textbf{\begin{tabular}[c]{@{}c@{}}Degree\\ Range\end{tabular}} \\ \hline
+ & $x_{min}$ & $\alpha$ & $x_{min}$ & $\mu$ &
+ $\sigma$ & \textit{Min-Max} \\ \hline
+\textit{User Degree} & 53 & 1.713442 & 2 & 2.963302 & 2.301555 & 1-18,374 \\ \hline
+\textit{Process Degree} & 1537 & 1.751258 & 74 & 7.798898 & 1.825076 & 1-1,030,145 \\ \hline
+\textit{File Degree} & 2 & 4.796094 & 1 & 0.7833566 & 0.3253502 & 2-1,371,165 \\ \hline
+\textit{Power-Law} & 7 & 1.952728 & 3 & -17.937673 & 4.869764 & 1-14,086 \\ \hline
+\end{tabular}
+\end{table*}
+
+\begin{figure}[h!]
+\begin{center}
+ \includegraphics[width=3in]{exps/user-histogram.pdf}
+ \caption{Histogram of User with different intervals.}
+ \label{f2}
+ \end{center}
+\end{figure}
+
+The first reason is that, for histogram figures, different intervals will significantly affect their shapes. In Fig.~\ref{f1}, we only show one interval case for each sample, it is far from enough to deny the hypothetical distribution. Fig.~\ref{f2} shows how the intervals of histogram affect the visual shapes of user node degrees. The figures with coarser granularity intervals clearly fit the power-law better as the two figures in the right column show. So, it will not be a good idea to exclude the power-law distribution just because one histogram figure does not look like power-law distribution.
+
+The second reason is even for the histogram figures in Fig.~\ref{f1}, the user and process entities can still be considered as obeying the power-law distribution if we consider the factor of $x_{min}$. It is not trivial to notice that both user and process figures have a straight tail, which shows that if we start from a specific $x$ instead of $1$, the distributions will be closer to power-law.
+
+\section{Using PoweRlaw Package}
+
+Another strategy to test the power-law hypothesis is to just consider the sample fits the power-law, then use \textit{maximum likelihood estimator} to estimate the possible parameters, which in power-law case are $x_{min}$ and $\alpha$. Using these parameter as an hypothesis, then we can use \textit{p-value} to quantify the plausibility of the hypothesis. This strategy was introduced in research work done by A. Clauset, et. al. There is also a package called \textit{poweRlaw} based on this strategy in R.
+
+So, based on this tool-chain, we look through the user, process, and file degree distribution again. For each sample, we also generate corresponding $x_{min}$ and $\alpha$. The results are shown as Fig.~\ref{f3} and Table~\ref{t1}.
+
+\begin{figure}[h!]
+\begin{center}
+ \includegraphics[width=3.2in]{exps/poweRlaw-cdf.pdf}
+ \caption{Degree CDF for different entities. The \textbf{red line} denotes Power-Law distribution, the \textbf{blue line} shows the Log-Normal distribution.}
+ \label{f3}
+ \end{center}
+\end{figure}
+
+Fig.~\ref{f3} contains four different subfigures. All of them are CDF figure based on doubly logarithm scale. The CDF figures are built based on this equation:
+\begin{equation}
+P(X) = \sum p(x) | x \leq X
+\end{equation}
+The benefit of using CDF instead of histograms is that we do not need to worry about the intervals any more.
+
+In the fourth subfigure of Fig.~\ref{f3}, we show a typical power-law distribution. It generates a straight line for an idea power-law distribution in the log-log scale. To show the possible fitted distributions, we also plot two extra lines, which represent the fitted \textit{power-law} distribution and \textit{log-normal} distribution with fitted parameters in all those four figures. The \textbf{red line} denotes power-law distribution, the \textbf{blue line} shows the log-normal distribution. The detailed fitted parameters are shown in Table 1.
+
+These results are pretty interesting if we compare them with Fig.~\ref{f1}. First, the user entities and process entities, which are not so close to power-law distribution in the histogram figures, now are proven to be quite fit to the power-law distribution. The $x_{min}$ values for users and processes are $53$ and $1537$ respectively from Table~\ref{t1}. Considering their wide degree ranges, the parameter values show they both are acceptable to be considered as power-law distributions.
+
+Second, the file entities turn out to be not fitting to power-law distribution so well as we expected. The poweRlaw tool chain returns a very bad estimation on both $x_{min}$ and $\alpha$. Although we can notice that the file degree CDF figure itself actually can fit a straight line well visually, the tool chain still gives a clearly not plausible estimation. The reason is not clear to me now, but it is highly possible that the proposed methodology has limitations on specific types of samples.
+
+Third, we may notice that both the log-normal and power-law distribution can fit the user, process, and even the power-law samples well. In some cases, the log-normal curves even are closer than power-law curves, like in the process degree figure. This actually tells us that, saying a data sample only fits the power-law distribution is still ambitious assertion due to current limited understanding on this complex distribution.
+
+\section{Conclusion}
+
+In this document, we introduce the basic concept of power-law distribution and show several ways to check whether given metadata samples fit the power-law distribution or not. Integrating the results from histogram figures and from PoweRlaw tool chain, we consider it would be appropriate to assert that the user, process, and file degrees fit the power-law distribution and we can generate synthetic metadata graph samples from power-law distribution. Note that, from the results, we can also observe that other distributions, like \textit{Log-Normal}, can also be used to describe these samples. However, we still consider power-law distribution is more favorable as it is proposed for complex network systems, which are more close to the context of our metadata graph.
+
+%\section{Future Work}
+
+\end{document}
\ No newline at end of file
diff --git a/papers/triton-wire-protocol/usenix.sty b/papers/meta-graph/report/usenix.sty
similarity index 100%
copy from papers/triton-wire-protocol/usenix.sty
copy to papers/meta-graph/report/usenix.sty
diff --git a/papers/meta-graph/rscripts/fdegree.r b/papers/meta-graph/rscripts/fdegree.r
index 7b66694..fc79d55 100644
--- a/papers/meta-graph/rscripts/fdegree.r
+++ b/papers/meta-graph/rscripts/fdegree.r
@@ -23,15 +23,17 @@ plot(log10(x2), log10(y4), type="l", col=4, lwd=3, yaxt="n", xaxt="n", bty="n",
xlim=c(0, 7), ylim=c(0, 8))
par(new=TRUE)
-plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Number of Vertices", xlim=c(0, 7), ylim=c(0, 8))
+#plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Number of Vertices", xlim=c(0, 7), ylim=c(0, 8))
+plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Vertices Number",
+ xlim=c(0, 7), ylim=c(0, 8), cex.lab=2)
ticks=seq(0, 7, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(1, at=c(0,1,2,3,4,5,6,7), labels=labels)
+axis(1, at=c(0,1,2,3,4,5,6,7), labels=labels, cex.axis=1.6)
ticks=seq(0, 8, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(2, at=c(0,1,2,3,4,5,6,7,8), labels=labels, las=1)
+axis(2, at=c(0,1,2,3,4,5,6,7,8), labels=labels, las=1, cex.axis=1.6)
legend("topright", legend = c(expression(paste(alpha, " = 4")),
expression(paste(alpha, " = 2")),
diff --git a/papers/meta-graph/rscripts/jdegree.r b/papers/meta-graph/rscripts/jdegree.r
index 1d27ead..62d9767 100644
--- a/papers/meta-graph/rscripts/jdegree.r
+++ b/papers/meta-graph/rscripts/jdegree.r
@@ -25,15 +25,16 @@ plot(log10(x2), log10(y4), type="l", col=4, lwd=3, yaxt="n", xaxt="n", bty="n",
xlim=c(0, maxX), ylim=c(0, maxY))
par(new=TRUE)
-plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Number of Vertices", xlim=c(0, maxX), ylim=c(0, maxY))
+plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Vertices Number",
+ xlim=c(0, maxX), ylim=c(0, maxY), cex.lab=2)
ticks=seq(0, maxX, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(1, at=ticks, labels=labels)
+axis(1, at=ticks, labels=labels, cex.axis=1.6)
ticks=seq(0, maxY, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(2, at=ticks, labels=labels, las=1)
+axis(2, at=ticks, labels=labels, las=1, cex.axis=1.6)
legend("topright", legend = c(expression(paste(alpha, " = 3")),
expression(paste(alpha, " = 1.7")),
diff --git a/papers/meta-graph/rscripts/pdegree.r b/papers/meta-graph/rscripts/pdegree.r
index 6607ce1..081acc0 100644
--- a/papers/meta-graph/rscripts/pdegree.r
+++ b/papers/meta-graph/rscripts/pdegree.r
@@ -25,15 +25,17 @@ plot(log10(x2), log10(y4), type="l", col=4, lwd=3, yaxt="n", xaxt="n", bty="n",
xlim=c(0, maxX), ylim=c(0, maxY))
par(new=TRUE)
-plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Number of Vertices", xlim=c(0, maxX), ylim=c(0, maxY))
+#plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Number of Vertices", xlim=c(0, maxX), ylim=c(0, maxY))
+plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Vertices Number",
+ xlim=c(0, maxX), ylim=c(0, maxY), cex.lab=2)
ticks=seq(0, maxX, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(1, at=ticks, labels=labels)
+axis(1, at=ticks, labels=labels, cex.axis=1.6)
ticks=seq(0, maxY, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(2, at=ticks, labels=labels, las=1)
+axis(2, at=ticks, labels=labels, las=1, cex.axis=1.6)
legend("topright", legend = c(expression(paste(alpha, " = 2")),
expression(paste(alpha, " = 1")),
diff --git a/papers/meta-graph/rscripts/udegree.r b/papers/meta-graph/rscripts/udegree.r
index e835922..e43cc0b 100644
--- a/papers/meta-graph/rscripts/udegree.r
+++ b/papers/meta-graph/rscripts/udegree.r
@@ -11,14 +11,15 @@ s = sum(y)
maxX=max(log10(x))
maxY=max(log10(y))
-plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Number of Vertices", xlim=c(0, maxX), ylim=c(0, maxY))
+plot(log10(x), log10(y), xaxt="n", yaxt="n", xlab="Degree", ylab="Vertices Number",
+ xlim=c(0, maxX), ylim=c(0, maxY), cex.lab=2)
ticks=seq(0, maxX, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(1, at=ticks, labels=labels)
+axis(1, at=ticks, labels=labels, cex.axis=1.6)
ticks=seq(0, maxY, by=1);
labels <- sapply(ticks, function(i) as.expression(bquote(10^ .(i))))
-axis(2, at=ticks, labels=labels, las=1)
+axis(2, at=ticks, labels=labels, las=1, cex.axis=1.6)
invisible(dev.off())
\ No newline at end of file
diff --git a/papers/meta-graph/usecases.tex b/papers/meta-graph/usecases.tex
index 88837ea..5312d3d 100644
--- a/papers/meta-graph/usecases.tex
+++ b/papers/meta-graph/usecases.tex
@@ -1,10 +1,10 @@
-\section{Use Cases On Using Metadata Graph}
+\section{Use Cases For the Metadata Graph}
-Unifying rich metadata into one graph turns many appealing data management functionalities into graph traversal operations or graph queries. In this section, we will show how to map the use cases from real-world scenarios to graph operations.
+Unifying rich metadata into one graph turns many appealing data management functionalities into graph traversal operations or graph queries. In this section, we will show how to map use cases from real-world scenarios to graph operations.
\subsection{User Audit}
-Data auditing is critical in large computing facilities where different users share the same cluster. It at least requires the detailed view of users file access for future security check. In metadata graph, we already collect the \textit{run} relationships between Users and Executions, and the \textit{read/write} relationships between Executions and Data Objects. And, all those relationships contain properties like timestamps. In such graph, the need to find all the files that were read by a specific user during given time frame [$t_s$, $t_e$] will become graph operations like this: 1) query the metadata graph from the given user; 2) travel through \textit{run} edges to Execution nodes; 3) filter executions based on the given time frame; 4) and travel through the \textit{read} edges to the final files. Similarly, if we want to get all the users who used to read to a sensitive file, we can do the similar graph query from the give file node.
+Data auditing is critical in large computing facilities where different users share the same cluster. It requires the detailed user-to-file access history for future security check. In metadata graph, we already collect the \textit{run} relationships between Users and Executions, and the \textit{read/write} relationships between Executions and Data Objects will record the file access history (all those relationships will contain properties like timestamps). Based on such graph, the need to find all the files that were read by a specific user during given time frame [$t_s$, $t_e$] will become graph operations like this: 1) locate the given user from the metadata graph; 2) travel through \textit{run} edges from this User node to Execution nodes; 3) filter executions based on the given time frame; 4) travel through the \textit{read} edges from the remained executions to the final files. Similarly, if we want to get all the users who used to read to a specific file, we can perfo
rm similar graph query from the give file node.
%This query can be expressed as a Gremlin script easily like following code shows:
%\begin{lstlisting}
@@ -23,7 +23,7 @@ Data auditing is critical in large computing facilities where different users sh
\subsection{Hierarchical Data Traversal}
-Hierarchical data organization is used to present a logical layout of data sets to users. The simplest example of hierarchical data traversal is traditional directory namespace traveling. In metadata graph model, we already abstract both the directories and files as Data Object entities. The \textit{belongs} and \textit{contains} relationships between different Data Objects represent the relationships between files and directories. So, given an absolute path, locating the file becomes going through a bunch of \textit{contains} edges from a Data Object node. Each time, we filter the edges according to the given names. Moreover, the access control metadata attached in users, files, and directories also could be verified while traversing.
+Hierarchical data organization is used to present a logical layout of data sets to users. The simplest example of hierarchical data traversal is traditional directory namespace traveling. In the metadata graph model, we abstract both directories and files as Data Object entities. The \textit{belongs} and \textit{contains} relationships between different Data Objects represent the relationships between files and directories. So, given an absolute path, locating a file becomes going through a bunch of \textit{contains} edges from a Data Object node. Each time, we filter the edges according to the given names. Moreover, the access control metadata attached in users, files, and directories also could be verified while traversing.
%\begin{lstlisting}
%// locate: /rootFS/dir/file.data
@@ -33,16 +33,16 @@ Hierarchical data organization is used to present a logical layout of data sets
% .filter{it.name = file.data}
%\end{lstlisting}
-An appealing advantage of using graph to store the directory structure of file system is the scalability. Traditional POSIX directory structure limits the number of files inside one directory. So, HPC system that may have millions of files under one directory, needs to deploy specific system like Giga+ to distribute the metadata into multiple servers for better performance~\cite{patil2011scale}. However, for proposed graph model, this turns into a graph partition problem, which we have been well studied~\cite{kim2012sbv, gonzalez2012powergraph, abou2006multilevel}.
+An appealing advantage of using a graph to store the directory structure is the scalability. Traditional POSIX directories have a limitation on the number of files inside one directory, so they can not provide sufficient performance for HPC systems where millions of files may be stored under one directory. To overcome this problem, research work like Giga+~\cite{patil2011scale} were proposed. The performance was improved by evenly distributing the metadata into multiple servers. However, in graph model, this turns out to be a graph partition problem, which already was well considered and studied~\cite{kim2012sbv, gonzalez2012powergraph, abou2006multilevel}.
-In addition to traditional POSIX-style files and directories, semantic data management would be another hierarchical traversal use case. Scientists usually need to manage their data in a semantic way, like arranging all of the inputs and outputs of a single simulation execution together. Traditionally, this needs careful file naming and directories placement. But in metadata graph, we can simply create new entity named Simulation and connect it with the Data Object with \textit{input/output} relationships. This will intelligently help users organize data in multiple dimensions.
+In addition to traditional POSIX-style files and directories, semantic data management would be another hierarchical traversal use case. Scientists usually need to manage their data in a semantic way, such as arranging all of the inputs and outputs of a single simulation execution together. Traditionally, this requires careful file naming and directory placement. But in a metadata graph, we can simply create new entity named Simulation and connect it with the Data Object. This help users organize data in multiple dimensions. %with \textit{input/output} relationships
\subsection{Provenance Support}
-Provenance has a wide range of use cases including data reproducibility, work-flow management etc. As a superset of provenance, the metadata graph model is able to support these usages. In this subsection, we borrow the problem from the first Provenance Challenge as an example~\cite{provchallengeweb}.
+Provenance has a wide range of use cases including data sharing, reproducibility, and work-flow management etc. As a superset of provenance, the metadata graph model naturally support these usages. In this subsection, we borrow the problem from the first Provenance Challenge as an example~\cite{provchallengeweb}.
-In this challenge, a simple example work-flow was provided as the basis, a workable provenance system should be able to represent the work-flow and all the relevant provenance for the example work-flow, and, most importantly, be able to answer predefined queries. Based on proposed graph model, we can easily abstract the work-flow as serial of executions run by the same user. Each execution reads several Data Objects and generates outputs for applications in next phase. Based on this work-flow, the provenance system needs to answer queries like: \textit{find the execution whose model is AlignWarp and inputs have annotation} [`center':`UChicago']. This can be expressed easily in the metadata graph: 1) query all the Execution vertices that have \textit{exe} out-edge pointing to a Data Object named `AlignWarp'; 2) start from all those Execution vertices and get those executions whose property `center' equal to `UChicago'.
+In this challenge, a simple example work-flow was provided as the basis. A workable provenance system should be able to represent the work-flow and all the relevant provenance, and, most importantly, be able to answer predefined queries. Based on our proposed graph model, it is straightforward to abstract the work-flow as series of executions run by the same user, and each execution reads files (Data Objects) and generates outputs for applications in the next phase. Based on this work-flow, the provenance system needs to answer queries like: \textit{find the execution whose model is AlignWarp and inputs have annotation} [`center':`UChicago']. This can be expressed easily in the metadata graph: 1) query all the Execution vertices, which have \textit{exe} out-edge pointing to a Data Object named `AlignWarp'; 2) start from all those Execution vertices and filter out the executions whose property `center' is not `UChicago'.
-A notable advantage of the metadata graph comparing with pure provenance system is that we can cross-reference different category of metadata in an unified way. If the provenance query needs the help of other metadata, like the file size, permission mode, or user group information etc., processing them in a unified graph will be more efficient and straightforwards.
+A notable advantage of the metadata graph compared to a pure provenance system is that it allows users to cross-reference different category of metadata in an unified way. If the provenance query needs the help of other metadata, like the file size, permission mode, or user group information, processing them in a unified graph will be more efficient and straightforward.
%. Here, we use the 8th query as an example:
hooks/post-receive
--
1
0
29 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 0363f97f87cf40e75f4ab55c4dd96c5ced648300 (commit)
from 819b08a17a4adf505d075d15e7eddc352611dc24 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 0363f97f87cf40e75f4ab55c4dd96c5ced648300
Author: John Jenkins <jenkins(a)mcs.anl.gov>
Date: Fri Aug 29 14:22:05 2014 -0500
hyperdex
-----------------------------------------------------------------------
Summary of changes:
notes/jenkins_related-data-analytics.pptx | Bin 129439 -> 131174 bytes
1 files changed, 0 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/notes/jenkins_related-data-analytics.pptx b/notes/jenkins_related-data-analytics.pptx
index 00fdf98..86b6b32 100644
Binary files a/notes/jenkins_related-data-analytics.pptx and b/notes/jenkins_related-data-analytics.pptx differ
hooks/post-receive
--
1
0
29 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 819b08a17a4adf505d075d15e7eddc352611dc24 (commit)
from 1f204a9e7b2040b253306afb629960cb0b7c7004 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 819b08a17a4adf505d075d15e7eddc352611dc24
Author: John Jenkins <jenkins(a)mcs.anl.gov>
Date: Fri Aug 29 08:59:22 2014 -0500
minor tweaks
-----------------------------------------------------------------------
Summary of changes:
presentations/ddn-2014/3-simulation.pptx | Bin 4011002 -> 4012467 bytes
1 files changed, 0 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/presentations/ddn-2014/3-simulation.pptx b/presentations/ddn-2014/3-simulation.pptx
index 58b789f..0661086 100644
Binary files a/presentations/ddn-2014/3-simulation.pptx and b/presentations/ddn-2014/3-simulation.pptx differ
hooks/post-receive
--
1
0
28 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 3ede42090115fb6eb8f6538e386e005b08333444 (commit)
from d294a6d04798e557a1a8357d543358ef9b65a255 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 3ede42090115fb6eb8f6538e386e005b08333444
Author: Kevin Harms <harms(a)alcf.anl.gov>
Date: Thu Aug 28 17:56:48 2014 -0500
Add script to convert automake tests to tap result file
-----------------------------------------------------------------------
Summary of changes:
code/scripts/jenkins/convert-trs-to-tap.py | 46 ++++++++++++++++++++++++++++
1 files changed, 46 insertions(+), 0 deletions(-)
create mode 100644 code/scripts/jenkins/convert-trs-to-tap.py
Diff of changes:
diff --git a/code/scripts/jenkins/convert-trs-to-tap.py b/code/scripts/jenkins/convert-trs-to-tap.py
new file mode 100644
index 0000000..b982d86
--- /dev/null
+++ b/code/scripts/jenkins/convert-trs-to-tap.py
@@ -0,0 +1,46 @@
+#!/usr/bin/env python
+#
+# Convert the .trs files generated by automake into a .tap file
+# for use by Jenkins test reporting.
+#
+
+import os
+import re
+
+def parse_trs(root, fname):
+ result = "not ok"
+ ro = re.compile(":test-result: (\w+)")
+ f = open(root+"/"+fname)
+ for line in f:
+ mo = ro.match(line)
+ if mo:
+ result = mo.group(1)
+ f.close()
+ return (fname, result)
+
+def gen_tap(f, outs):
+ count = 1
+ print >>f, "1..%d"%(len(outs))
+ for out in outs:
+ if out[1] == "PASS":
+ print >>f, "ok ", count, " - ", out[0]
+ elif out[1] == "XFAIL":
+ print >>f, "not ok ", count, " - ", out[0], "# expected failure"
+ else:
+ print >>f, "not ok ", count, " - ", out[0]
+ count += 1
+ return
+
+def main():
+ outs = []
+ for root, dirs, files in os.walk("build/tritonbuild/tests"):
+ for file in files:
+ if file.endswith('.trs'):
+ outs.append(parse_trs(root, file))
+ f = open("build/tritonbuild/tests/results.tap", "w")
+ gen_tap(f, outs)
+ f.close()
+ return
+
+if __name__ == '__main__':
+ main()
hooks/post-receive
--
1
0
28 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 1f204a9e7b2040b253306afb629960cb0b7c7004 (commit)
from 10adc1a1517cfe03336aa65a005361c5a89f076f (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 1f204a9e7b2040b253306afb629960cb0b7c7004
Author: Rob Ross <rbross(a)mailinator.com>
Date: Thu Aug 28 16:09:50 2014 -0500
added Voulgaris thesis notes.
-----------------------------------------------------------------------
Summary of changes:
notes/ross_related-data-analytics.pptx | Bin 2939172 -> 2940952 bytes
1 files changed, 0 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/notes/ross_related-data-analytics.pptx b/notes/ross_related-data-analytics.pptx
index 5cceae6..7cfb468 100644
Binary files a/notes/ross_related-data-analytics.pptx and b/notes/ross_related-data-analytics.pptx differ
hooks/post-receive
--
1
0
28 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 10adc1a1517cfe03336aa65a005361c5a89f076f (commit)
via 8f980ff07951626dfaa4da4376659b6847860d76 (commit)
from 8537d7cd967922b4b51819e82a2b29f06085dfba (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 10adc1a1517cfe03336aa65a005361c5a89f076f
Merge: 8f980ff07951626dfaa4da4376659b6847860d76 8537d7cd967922b4b51819e82a2b29f06085dfba
Author: Rob Ross <rbross(a)mailinator.com>
Date: Thu Aug 28 15:08:47 2014 -0500
Merge branch 'master' of git.mcs.anl.gov:triton-private
commit 8f980ff07951626dfaa4da4376659b6847860d76
Author: Rob Ross <rbross(a)mailinator.com>
Date: Thu Aug 28 15:08:30 2014 -0500
bunch of notes
-----------------------------------------------------------------------
Summary of changes:
notes/ross_related-data-analytics.pptx | Bin 2751568 -> 2939172 bytes
1 files changed, 0 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/notes/ross_related-data-analytics.pptx b/notes/ross_related-data-analytics.pptx
index 511aa47..5cceae6 100644
Binary files a/notes/ross_related-data-analytics.pptx and b/notes/ross_related-data-analytics.pptx differ
hooks/post-receive
--
1
0
26 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 8537d7cd967922b4b51819e82a2b29f06085dfba (commit)
from 2d9f8d916cdf8daf9695150384d3feff4724d297 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 8537d7cd967922b4b51819e82a2b29f06085dfba
Author: Phil Carns <carns(a)mcs.anl.gov>
Date: Tue Aug 26 15:39:53 2014 -0400
minor updates
-----------------------------------------------------------------------
Summary of changes:
presentations/ddn-2014/2-triton.pptx | Bin 599909 -> 599942 bytes
1 files changed, 0 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/presentations/ddn-2014/2-triton.pptx b/presentations/ddn-2014/2-triton.pptx
index 20e0f64..86a4044 100644
Binary files a/presentations/ddn-2014/2-triton.pptx and b/presentations/ddn-2014/2-triton.pptx differ
hooks/post-receive
--
1
0
25 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 2d9f8d916cdf8daf9695150384d3feff4724d297 (commit)
from 48dab9dc89c60e1fb0db786aa8474c229b01637f (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 2d9f8d916cdf8daf9695150384d3feff4724d297
Author: John Jenkins <jenkins(a)mcs.anl.gov>
Date: Mon Aug 25 16:09:20 2014 -0500
pregel, giraph info
-----------------------------------------------------------------------
Summary of changes:
notes/jenkins_related-data-analytics.pptx | Bin 130593 -> 129439 bytes
1 files changed, 0 insertions(+), 0 deletions(-)
Diff of changes:
diff --git a/notes/jenkins_related-data-analytics.pptx b/notes/jenkins_related-data-analytics.pptx
index 937574c..00fdf98 100644
Binary files a/notes/jenkins_related-data-analytics.pptx and b/notes/jenkins_related-data-analytics.pptx differ
hooks/post-receive
--
1
0
25 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via 48dab9dc89c60e1fb0db786aa8474c229b01637f (commit)
via 1ece93856662e8732ab97eb635af76022f2baad6 (commit)
via 1f78fcfa3952a398555e3cfc3760d07c248df0cd (commit)
from be45fabab1155e660e5b808f74563d9ea07a53fb (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit 48dab9dc89c60e1fb0db786aa8474c229b01637f
Merge: 1ece93856662e8732ab97eb635af76022f2baad6 be45fabab1155e660e5b808f74563d9ea07a53fb
Author: Rob Ross <rbross(a)mailinator.com>
Date: Mon Aug 25 10:17:47 2014 -0500
Merge branch 'master' of git.mcs.anl.gov:triton-private
commit 1ece93856662e8732ab97eb635af76022f2baad6
Author: Rob Ross <rbross(a)mailinator.com>
Date: Mon Aug 25 10:16:40 2014 -0500
removed this file.
commit 1f78fcfa3952a398555e3cfc3760d07c248df0cd
Author: Rob Ross <rbross(a)mailinator.com>
Date: Mon Aug 25 10:15:42 2014 -0500
adding gitignore.
-----------------------------------------------------------------------
Summary of changes:
papers/{swim => meta-graph}/.gitignore | 0
1 files changed, 0 insertions(+), 0 deletions(-)
copy papers/{swim => meta-graph}/.gitignore (100%)
Diff of changes:
diff --git a/papers/swim/.gitignore b/papers/meta-graph/.gitignore
similarity index 100%
copy from papers/swim/.gitignore
copy to papers/meta-graph/.gitignore
hooks/post-receive
--
1
0
25 Aug '14
This is an automated email from the git hooks/post-receive script. It was
generated because a ref change was pushed to the repository containing
the project "".
The branch, master has been updated
via d294a6d04798e557a1a8357d543358ef9b65a255 (commit)
from 35579e85e96c04af2b0e8c742f5539b62c476c91 (commit)
Those revisions listed above that are new to this repository have
not appeared on any other notification email; so we list those
revisions in full, below.
- Log -----------------------------------------------------------------
commit d294a6d04798e557a1a8357d543358ef9b65a255
Author: Kevin Harms <harms(a)alcf.anl.gov>
Date: Mon Aug 25 09:49:20 2014 -0500
jenkins: removed -fno-diagnostics-show-caret cflags which isn't supported on older compiler versions
-----------------------------------------------------------------------
Summary of changes:
code/scripts/jenkins/build.sh | 3 ++-
1 files changed, 2 insertions(+), 1 deletions(-)
Diff of changes:
diff --git a/code/scripts/jenkins/build.sh b/code/scripts/jenkins/build.sh
index b52c391..811afa1 100755
--- a/code/scripts/jenkins/build.sh
+++ b/code/scripts/jenkins/build.sh
@@ -200,7 +200,8 @@ cd "${SRCDIR}/code"
export PKG_CONFIG_PATH=${DEPINSTALL}/lib/pkgconfig:$PKG_CONFIG_PATH
echo "* Configuring Triton"
-CONFIGOPTS="--prefix=${INSTALLDIR} CFLAGS=-fno-diagnostics-show-caret"
+#CONFIGOPTS="--prefix=${INSTALLDIR} CFLAGS=-fno-diagnostics-show-caret"
+CONFIGOPTS="--prefix=${INSTALLDIR}"
cd "${BUILDDIR}"
"${SRCDIR}/code/configure" ${CONFIGOPTS} || exit 3
hooks/post-receive
--
1
0