Mg-rast
Threads by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2018 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2017 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2016 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2015 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2014 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2013 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2012 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2011 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2010 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2009 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2008 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2007 -----
- December
- November
- October
- September
- August
- July
July 2011
- 55 participants
- 221 discussions
Hello there William,
That solved our doubts. Thanks for the help. But that means we won't be able to do the analyses we wanted. :/
Anyway, if that's the case I would like to offer a suggestion. You see, one of the advantages of metagenomic analysis is the possibility of analysing both taxonomic composition and function. One of the main goals of microbial ecology is the crossing of both these informations. So I think you are missing a great opportunity to offer a tool for this comparison. I did so previously in MG-Rast v2 by separating a given subset of sequences and re-submitting them. If something like the workbench could offer valid counts for this sort of crossed information it would be great and more practical than the approach I took. That would allow the exploration of functions associated to a given taxonomic group, or taxonomic composition of a particular subsystem, opening up whole new possibilities of analysis. My data revealed trends unnoticed when the information was crossed than when the analyses were separated.
So... thanks again,
Best regards!
Gustavo
Gustavo Bueno Gregoracci
--
MSc., Dr. in Microbiology (Genetics and Molecular Biology)
Email: gustavo_biomed(a)yahoo.com
--
Pos doc Student
Laboratório de Microbiologia (Prof. Fabiano Thompson)
Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
Tel: +55 (21) 2562-6567
>________________________________
>From: William Trimble <wltrimbl(a)gmail.com>
>To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>; mg-rast(a)mcs.anl.gov
>Sent: Friday, July 29, 2011 5:19 PM
>Subject: Re: [mg-rast] Doubt about MG-Rast annotation
>
>> ...
>> Generate a Domain Level table.
>>
>> Now we have only two domains, which will make it easier to explain. Mark had
>> explained us the “abundance” and “hits” fields like you did. We understood
>> it when they refer to a metagenome, but they make less sense when working
>> with the Workbench.
>I think I understand what's going on.
>
>When the organism and function tables are generated, there is a filter that
>chooses the best similarity hit for each protein query. The "abundance"
>numbers in the organism and function tables are intended to reflect the
>abundance of a function or organism.
>
>On the workbench, the database hits are not subjected to the same
>"best-hit-only" filter. Multiple annotations can be associated with each
>sequence. This results in some of the sequences on the workbench appearing
>multiple times with different annotations, both functional or taxonomic.
>
>The workbench was designed to retrieve sets of sequences, not to count them.
>It is for this reason the numbers don't add up; some of the sequences hit
>proteins in the database that had hundreds of annotations.
>
>I hope this helps.
>
>William Trimble
>--for the MG-RAST team
>
>> We imagined that when you select things to send to the workbench and check
>> the “use proteins from workbench” option, the analysis would be restricted
>> to the set of things sent to the workbench, in this case to the 74 unique
>> proteins and 221 total hits. So we get 56 bacterial proteins and 101
>> eukaryotic ones, adding up to 157, instead of the total 221 sent. But this
>> loss is understandable since the databases are different and possibly not
>> all 221 are identified in the other database. First thing is to know if this
>> interpretation is right.
>>
>> Second doubt is about the “#hits” in this domain table, or the total number
>> of unique sequences generated from the workbench. If our above explanation
>> is correct, we suppose the analysis would be restricted to the 221 hits. How
>> do they turn into the new “#hits” of this new table? Let’s get to the
>> numbers… From the 221 hits sent to the workbench, 157 were identified
>> (“workbench abundance” bacteria + eukarya). How this becomes 97 unique
>> proteins for bacteria and 163 for eukarya (“#hits”)? 163+97 add up to the
>> 260 which is more than the total of hits sent to the workbench (221) and the
>> total identified (157).
>>
>> Similarly, how the new “abundance” is generated? If we sent 221 sequences to
>> the workbench, and 157 were identified, where do 245 bacterial and 403
>> eukaryotic sequences come from? They add up to 648!
>>
>> Sorry to bother you,
>> Thanks for the help!!
>> Best regards!
>> Gustavo
>>
>> Gustavo Bueno Gregoracci
>> --
>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> Email: gustavo_biomed(a)yahoo.com
>> --
>> Pos doc Student
>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> Tel: +55 (21) 2562-6567
>>
>> ________________________________
>> From: William Trimble <wltrimbl(a)gmail.com>
>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>; mg-rast(a)mcs.anl.gov
>> Sent: Friday, July 29, 2011 2:32 PM
>> Subject: Re: [mg-rast] Doubt about MG-Rast annotation
>>
>> Sorry for the delay in getting back to you.
>>
>>> Let's use an example, like the metagenome I sent you. We used the
>>> Subsytems
>>> database and separated Photosynthesis with 209 of abundance and 66 unique
>>> proteins in the group. We sent it to the Workbench to look into the
>>> taxonomic profile (Genbank).
>>
>> This means that there were 66 distinct sequences in the database that were
>> hit by a total of 209 predicted genes in your sequence data.
>>
>>> 1. It stored all the 209 hits from the abundance in the workbench, right?
>>> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
>>> Bacteria, which is shown in the workbench abundance column. Is this
>>> reasoning right? This loss is the result of the conversion between
>>> databases?
>>> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
>>> each sequence generate more than one annotation? Cause we got for bacteria
>>> an abundance of 245 and for eukarya an abundance of 403, from the 209
>>> sequences selected in the workbench.
>>> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
>>> from the 209 hits (see question 1), how do they turn into 97 unique
>>> proteins
>>> for bacteria and 163 unique proteins for eukarya?
>> This is a bit of a puzzle.
>>
>> Let me make sure we're doing exactly the same thing.
>> Go to analysis page, Functional annotation tab
>> Select 4465448.3 with the default annotation source subsystems
>> Generate a table
>> group by level 1
>> select photosynthesis, with
>> abundance=221 and proteins=74
>> click To workbench.
>> What are your next steps?
>>
>> The "abundance" field is intended to be proportional to the number of times
>> a
>> given annotation is present in the data; the "hits" only counts unique
>> sequences in the database.
>>
>> I'm not immediately sure why subsets of the sequences on the workbench
>> have unexpected sizes.
>>
>> William Trimble
>> --for the MG-RAST team
>>
>>> Again, I thank you for the time and patience,
>>> Looking forward to hear from you,
>>> Best regards,
>>> Gustavo
>>>
>>> Gustavo Bueno Gregoracci
>>> --
>>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> Email: gustavo_biomed(a)yahoo.com
>>> --
>>> Pos doc Student
>>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> Tel: +55 (21) 2562-6567
>>>
>>> ________________________________
>>> From: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>>> To: Mark DSouza <dsouza(a)mcs.anl.gov>
>>> Cc: Louisi Oliveira <louisioliveira(a)gmail.com>; MG- Rast
>>> <mg-rast(a)mcs.anl.gov>
>>> Sent: Wednesday, July 20, 2011 5:53 PM
>>> Subject: Re: Doubt about MG-Rast annotation
>>>
>>> Hi there,
>>> Thanks for the help so far. We cleared some important questions but there
>>> are new ones.
>>>
>>> Let's use an example, like the metagenome I sent you. We used the
>>> Subsytems
>>> database and separated Photosynthesis with 209 of abundance and 66 unique
>>> proteins in the group. We sent it to the Workbench to look into the
>>> taxonomic profile (Genbank).
>>> 1. It stored all the 209 hits from the abundance in the workbench, right?
>>> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
>>> Bacteria, which is shown in the workbench abundance column. Is this
>>> reasoning right? This is the conversion between databases?
>>>
>>> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
>>> each sequence generate more than one annotation? Cause we got for bacteria
>>> an abundance of 245 and for eukarya an abundance of 403, from the 209
>>> sequences selected in the workbench.
>>> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
>>> from the 209 hits (see question 1), how do they turn into 97 unique
>>> proteins
>>> for bacteria and 163 unique proteins for eukarya?
>>>
>>> Sorry to trouble you again, but we really need to clarify these questions
>>> to
>>> continue with our analyses. We also will be spreading this information
>>> among
>>> our colleagues.
>>>
>>> Thanks again!
>>> Best regards!
>>> Gustavo
>>>
>>> Gustavo Bueno Gregoracci
>>> --
>>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> Email: gustavo_biomed(a)yahoo.com
>>> --
>>> Pos doc Student
>>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> Tel: +55 (21) 2562-6567
>>>
>>> ________________________________
>>> From: Mark DSouza <dsouza(a)mcs.anl.gov>
>>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>>> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
>>> <louisioliveira(a)gmail.com>
>>> Sent: Wednesday, July 20, 2011 2:44 PM
>>> Subject: Re: Doubt about MG-Rast annotation
>>>
>>> Hi,
>>>
>>>> The "# proteins" field gives me the number of unique sequences which
>>>> matched the genbank database ...
>>>
>>> No, this is not accurate, the '# proteins' field gives the number of
>>> unique
>>> sequences within the selected grouping so if you group by phylum you will
>>> get unique sequences within the phylum.
>>>
>>>> ... the "abundance" is the total number of matches including repeated
>>>> matches to the "# proteins" proteins. Is that so?
>>>
>>> Yes, this is correct.
>>>
>>>> When I send these to the workbench the "# proteins" is sent; and this
>>>> subset is re-annotated with Subsystems, for example. The "#proteins"
>>>> falls
>>>> since not all of the original set annotated with Genbank will be
>>>> recognized
>>>> here, right?
>>>
>>> That is right, there may be sequences in GenBank which are not present in
>>> the SEED Subsystems database and vice versa.
>>>
>>>> But the abundance may increase because more repeats can be found in the
>>>> larger database?
>>>
>>> When moving to a database with a larger number of sequences, the abundance
>>> will be expected to increase because of the greater chance of finding a
>>> match.
>>>
>>>> And what the "workbench abundance" means?
>>> The workbench abundance is the count of the reads in the metagenomic
>>> sample
>>> with hits against the protein sequences (from database A, e.g. Subsystems)
>>> selected in the workbench which were found in the database B (e.g.
>>> GenBank).
>>> In this example database A is used to select the proteins placed in the
>>> workbench and database B is used to create the organism classification,
>>> using the proteins from the workbench.
>>>
>>> Regards,
>>> Mark
>>>
>>>
>>> ----- Original Message -----
>>>> From: "Gustavo B. Gregoracci" <gustavo_biomed(a)yahoo.com>
>>>> To: "Mark DSouza" <dsouza(a)mcs.anl.gov>
>>>> Cc: "MG- Rast" <mg-rast(a)mcs.anl.gov>, "Louisi Oliveira"
>>>> <louisioliveira(a)gmail.com>
>>>> Sent: Tuesday, July 19, 2011 8:46:22 PM
>>>> Subject: Re: Doubt about MG-Rast annotation
>>>> Hi there!
>>>>
>>>>
>>>> I'm not sure I got it right. Let's see... I annotated the metagenome
>>>> through Genbank, for example. The "# proteins" field gives me the
>>>> number of unique sequences which matched the genbank database and the
>>>> "abundance" is the total number of matches including repeated matches
>>>> to the "# proteins" proteins. Is that so? When I send these to the
>>>> workbench the "# proteins" is sent; and this subset is re-annotated
>>>> with Subsystems, for example. The "#proteins" falls since not all of
>>>> the original set annotated with Genbank will be recognized here,
>>>> right? But the abundance may increase because more repeats can be
>>>> found in the larger database?
>>>>
>>>>
>>>>
>>>> And what the "workbench abundance" means?
>>>>
>>>>
>>>> Sorry to bother you further, but I'm could not grasp the explanation
>>>> previously,
>>>> Thanks again for the help and patience,
>>>>
>>>>
>>>> Best regards,
>>>> Gustavo
>>>>
>>>>
>>>> Gustavo Bueno Gregoracci
>>>> --
>>>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>>> Email: gustavo_biomed(a)yahoo.com
>>>>
>>>> --
>>>> Pos doc Student
>>>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>>> Tel: +55 (21) 2562-6567
>>>>
>>>>
>>>>
>>>>
>>>>
>>>> From: Mark DSouza <dsouza(a)mcs.anl.gov>
>>>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>>>> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
>>>> <louisioliveira(a)gmail.com>
>>>> Sent: Monday, July 18, 2011 7:13 PM
>>>> Subject: Re: Doubt about MG-Rast annotation
>>>>
>>>> Hi,
>>>>
>>>> The number discrepancy arises from the use of different databases, the
>>>> original 209 hits is specific to the SEED subsystem, with 66 hits
>>>> against the proteins in these organisms. When you perform the organism
>>>> classification against GenBank, the number of hits is higher because
>>>> of the greater number of organisms in the GenBank database, and so
>>>> does the abundance counts. The workbench abundance column of the table
>>>> gives a better breakdown of the taxonomic counts as compared to the
>>>> 209 count. This is not very intuitive from the display and we are
>>>> looking into changing the table to make this more explicit.
>>>>
>>>> Regards,
>>>> Mark
>>>>
>>>> ----- Original Message -----
>>>> > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>>>> > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>>>> > Cc: "MG- Rast" < mg-rast(a)mcs.anl.gov >, "Louisi Oliveira" <
>>>> > louisioliveira(a)gmail.com >
>>>> > Sent: Monday, July 18, 2011 1:03:28 PM
>>>> > Subject: Re: Doubt about MG-Rast annotation
>>>> > Hi there Mark,
>>>> >
>>>> >
>>>> > Thanks for the reply. So, the metagenome ID is 4465448.3 and its
>>>> > called Forno. The glitch has actually happened in others as well,
>>>> > and
>>>> > I'm only using this one as an example.
>>>> >
>>>> >
>>>> > I re-checked the numbers and they still don't add. I'm working with
>>>> > Genbank (e-5) for organism classification, and subsystems (e-5) for
>>>> > functional classification, for that matter.
>>>> >
>>>> >
>>>> > I have 209 hits in this entire metagenome for photosynthesis. If I
>>>> > check the number of photosynthesis hits among the bacterial hits I
>>>> > get
>>>> > 140, while the same for eukaryotic gives me 142 hits. That puzzled
>>>> > me.
>>>> > When I do the opposite, and check the taxonomic composition of just
>>>> > the 209 photosynthesis hits the numbers are even more weirder. I get
>>>> > 245 hits for bacteria and 403 for eukaryotes!
>>>> >
>>>> >
>>>> > Thanks again for the help,
>>>> > I'm available for any further explanations necessary,
>>>> >
>>>> >
>>>> > Best regards,
>>>> > Gustavo
>>>> >
>>>> > Gustavo Bueno Gregoracci
>>>> > --
>>>> > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>>> > Email: gustavo_biomed(a)yahoo.com
>>>> >
>>>> > --
>>>> > Pos doc Student
>>>> > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>>> > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>>> > Tel: +55 (21) 2562-6567
>>>> >
>>>> >
>>>> >
>>>> >
>>>> >
>>>> > From: Mark DSouza < dsouza(a)mcs.anl.gov >
>>>> > To: Gustavo B. Gregoracci < gustavo_biomed(a)yahoo.com >
>>>> > Cc: MG- Rast < mg-rast(a)mcs.anl.gov >
>>>> > Sent: Monday, July 18, 2011 1:33 PM
>>>> > Subject: Re: Doubt about MG-Rast annotation
>>>> >
>>>> > Hi,
>>>> >
>>>> > Can you please send us the MG-RAST ID for the dataset where you
>>>> > observed this, we will check into it for you.
>>>> >
>>>> > --
>>>> > Regards,
>>>> > Mark D'Souza
>>>> > -- for the MG-RAST team
>>>> >
>>>> > mg-rast(a)mcs.anl.gov
>>>> > http://metagenomics.anl.gov/
>>>> >
>>>> > If you need to respond to this email please address it to the
>>>> > mailing-list mg-rast(a)mcs.anl.gov
>>>> >
>>>> >
>>>> > ----- Original Message -----
>>>> > > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>>>> > > To: "Mark D'Souza" < dsouza(a)mcs.anl.gov >
>>>> > > Sent: Saturday, July 16, 2011 10:05:14 AM
>>>> > > Subject: Doubt about MG-Rast annotation
>>>> > > Hello there Mark,
>>>> > >
>>>> > >
>>>> > >
>>>> > > As I was analyzing some metagenomes, I ran into a recurrent
>>>> > > problem.
>>>> > > I
>>>> > > was trying to cross information from taxonomic and functional
>>>> > > annotation, and discovered some inconsistencies. I could not
>>>> > > figure
>>>> > > those out so I’m writing you in the hope that you can help me,
>>>> > > since
>>>> > > this will affect my analysis.
>>>> > >
>>>> > >
>>>> > > You see… I wanted to understand the contribution of different
>>>> > > phylogenetic groups to a given subsystem. So I got the total
>>>> > > number
>>>> > > of
>>>> > > sequences for the photosynthesis subsystem, for example, which was
>>>> > > 209
>>>> > > hits. Then I went into taxonomy and separated all entries
>>>> > > identified
>>>> > > as eukaryotes, sending them to the workbench. I changed again to
>>>> > > functional and checked the total number of eukaryotic hits to
>>>> > > photosynthesis, which was 163. So far so good, but I decided to
>>>> > > perform the same analysis regarding prokaryotes. Sent all
>>>> > > prokaryotic
>>>> > > entries to the workbench and changed to functional to see their
>>>> > > contribution to photosynthesis as well. I got 174 hits!
>>>> > >
>>>> > >
>>>> > >
>>>> > > How is this possible? If my whole metagenome has 209 hits to a
>>>> > > subsystem, how can I have 163 eukaryotic hits within it and 174
>>>> > > prokaryotic hits to the same subsystem? Can sequences be annotated
>>>> > > as
>>>> > > both eukaryotic and prokaryotic? Cause I was interpreting these
>>>> > > categories as mutually exclusive…
>>>> > >
>>>> > >
>>>> > > Thanks in advance for the help,
>>>> > > Looking forward to hear from you,
>>>> > >
>>>> > >
>>>> > > Best regards,
>>>> > > Gustavo
>>>> > >
>>>> > > Gustavo Bueno Gregoracci
>>>> > > --
>>>> > > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>>> > > Email: gustavo_biomed(a)yahoo.com
>>>> > >
>>>> > > --
>>>> > > Pos doc Student
>>>> > > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>>> > > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>>> > > Tel: +55 (21) 2562-6567
>>>
>>>
>>>
>>>
>>>
>>
>>
>>
>
>
>
1
0
---------- Forwarded message ----------
From: William Trimble <wltrimbl(a)gmail.com>
Date: Fri, Jul 29, 2011 at 4:10 PM
Subject: Re: [mg-rast] Greengenes Database version
To: Gregory Ditzler <gregory.ditzler(a)gmail.com>
Thanks for using MG-RAST.
The database that we use is described on
http://metagenomics.anl.gov/metagenomics.cgi?page=Sources
It states that the current greengeenes version dates from
February 2011.
William Trimble
--for the MG-RAST Team
On Fri, Jul 29, 2011 at 3:49 PM, Gregory Ditzler
<gregory.ditzler(a)gmail.com> wrote:
> I am currently using your server and have a quick question about the
> analysis tools. What is the version of Greengenes that is currently being
> used by MG-RAST?
> Kind Regards
> Gregory Ditzler
> gregory.ditzler(a)gmail.com
1
0
I am currently using your server and have a quick question about the
analysis tools. What is the version of Greengenes that is currently being
used by MG-RAST?
Kind Regards
Gregory Ditzler
gregory.ditzler(a)gmail.com
1
0
> ...
> Generate a Domain Level table.
>
> Now we have only two domains, which will make it easier to explain. Mark had
> explained us the “abundance” and “hits” fields like you did. We understood
> it when they refer to a metagenome, but they make less sense when working
> with the Workbench.
I think I understand what's going on.
When the organism and function tables are generated, there is a filter that
chooses the best similarity hit for each protein query. The "abundance"
numbers in the organism and function tables are intended to reflect the
abundance of a function or organism.
On the workbench, the database hits are not subjected to the same
"best-hit-only" filter. Multiple annotations can be associated with each
sequence. This results in some of the sequences on the workbench appearing
multiple times with different annotations, both functional or taxonomic.
The workbench was designed to retrieve sets of sequences, not to count them.
It is for this reason the numbers don't add up; some of the sequences hit
proteins in the database that had hundreds of annotations.
I hope this helps.
William Trimble
--for the MG-RAST team
> We imagined that when you select things to send to the workbench and check
> the “use proteins from workbench” option, the analysis would be restricted
> to the set of things sent to the workbench, in this case to the 74 unique
> proteins and 221 total hits. So we get 56 bacterial proteins and 101
> eukaryotic ones, adding up to 157, instead of the total 221 sent. But this
> loss is understandable since the databases are different and possibly not
> all 221 are identified in the other database. First thing is to know if this
> interpretation is right.
>
> Second doubt is about the “#hits” in this domain table, or the total number
> of unique sequences generated from the workbench. If our above explanation
> is correct, we suppose the analysis would be restricted to the 221 hits. How
> do they turn into the new “#hits” of this new table? Let’s get to the
> numbers… From the 221 hits sent to the workbench, 157 were identified
> (“workbench abundance” bacteria + eukarya). How this becomes 97 unique
> proteins for bacteria and 163 for eukarya (“#hits”)? 163+97 add up to the
> 260 which is more than the total of hits sent to the workbench (221) and the
> total identified (157).
>
> Similarly, how the new “abundance” is generated? If we sent 221 sequences to
> the workbench, and 157 were identified, where do 245 bacterial and 403
> eukaryotic sequences come from? They add up to 648!
>
> Sorry to bother you,
> Thanks for the help!!
> Best regards!
> Gustavo
>
> Gustavo Bueno Gregoracci
> --
> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
> Email: gustavo_biomed(a)yahoo.com
> --
> Pos doc Student
> Laboratório de Microbiologia (Prof. Fabiano Thompson)
> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
> Tel: +55 (21) 2562-6567
>
> ________________________________
> From: William Trimble <wltrimbl(a)gmail.com>
> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>; mg-rast(a)mcs.anl.gov
> Sent: Friday, July 29, 2011 2:32 PM
> Subject: Re: [mg-rast] Doubt about MG-Rast annotation
>
> Sorry for the delay in getting back to you.
>
>> Let's use an example, like the metagenome I sent you. We used the
>> Subsytems
>> database and separated Photosynthesis with 209 of abundance and 66 unique
>> proteins in the group. We sent it to the Workbench to look into the
>> taxonomic profile (Genbank).
>
> This means that there were 66 distinct sequences in the database that were
> hit by a total of 209 predicted genes in your sequence data.
>
>> 1. It stored all the 209 hits from the abundance in the workbench, right?
>> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
>> Bacteria, which is shown in the workbench abundance column. Is this
>> reasoning right? This loss is the result of the conversion between
>> databases?
>> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
>> each sequence generate more than one annotation? Cause we got for bacteria
>> an abundance of 245 and for eukarya an abundance of 403, from the 209
>> sequences selected in the workbench.
>> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
>> from the 209 hits (see question 1), how do they turn into 97 unique
>> proteins
>> for bacteria and 163 unique proteins for eukarya?
> This is a bit of a puzzle.
>
> Let me make sure we're doing exactly the same thing.
> Go to analysis page, Functional annotation tab
> Select 4465448.3 with the default annotation source subsystems
> Generate a table
> group by level 1
> select photosynthesis, with
> abundance=221 and proteins=74
> click To workbench.
> What are your next steps?
>
> The "abundance" field is intended to be proportional to the number of times
> a
> given annotation is present in the data; the "hits" only counts unique
> sequences in the database.
>
> I'm not immediately sure why subsets of the sequences on the workbench
> have unexpected sizes.
>
> William Trimble
> --for the MG-RAST team
>
>> Again, I thank you for the time and patience,
>> Looking forward to hear from you,
>> Best regards,
>> Gustavo
>>
>> Gustavo Bueno Gregoracci
>> --
>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> Email: gustavo_biomed(a)yahoo.com
>> --
>> Pos doc Student
>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> Tel: +55 (21) 2562-6567
>>
>> ________________________________
>> From: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>> To: Mark DSouza <dsouza(a)mcs.anl.gov>
>> Cc: Louisi Oliveira <louisioliveira(a)gmail.com>; MG- Rast
>> <mg-rast(a)mcs.anl.gov>
>> Sent: Wednesday, July 20, 2011 5:53 PM
>> Subject: Re: Doubt about MG-Rast annotation
>>
>> Hi there,
>> Thanks for the help so far. We cleared some important questions but there
>> are new ones.
>>
>> Let's use an example, like the metagenome I sent you. We used the
>> Subsytems
>> database and separated Photosynthesis with 209 of abundance and 66 unique
>> proteins in the group. We sent it to the Workbench to look into the
>> taxonomic profile (Genbank).
>> 1. It stored all the 209 hits from the abundance in the workbench, right?
>> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
>> Bacteria, which is shown in the workbench abundance column. Is this
>> reasoning right? This is the conversion between databases?
>>
>> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
>> each sequence generate more than one annotation? Cause we got for bacteria
>> an abundance of 245 and for eukarya an abundance of 403, from the 209
>> sequences selected in the workbench.
>> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
>> from the 209 hits (see question 1), how do they turn into 97 unique
>> proteins
>> for bacteria and 163 unique proteins for eukarya?
>>
>> Sorry to trouble you again, but we really need to clarify these questions
>> to
>> continue with our analyses. We also will be spreading this information
>> among
>> our colleagues.
>>
>> Thanks again!
>> Best regards!
>> Gustavo
>>
>> Gustavo Bueno Gregoracci
>> --
>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> Email: gustavo_biomed(a)yahoo.com
>> --
>> Pos doc Student
>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> Tel: +55 (21) 2562-6567
>>
>> ________________________________
>> From: Mark DSouza <dsouza(a)mcs.anl.gov>
>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
>> <louisioliveira(a)gmail.com>
>> Sent: Wednesday, July 20, 2011 2:44 PM
>> Subject: Re: Doubt about MG-Rast annotation
>>
>> Hi,
>>
>>> The "# proteins" field gives me the number of unique sequences which
>>> matched the genbank database ...
>>
>> No, this is not accurate, the '# proteins' field gives the number of
>> unique
>> sequences within the selected grouping so if you group by phylum you will
>> get unique sequences within the phylum.
>>
>>> ... the "abundance" is the total number of matches including repeated
>>> matches to the "# proteins" proteins. Is that so?
>>
>> Yes, this is correct.
>>
>>> When I send these to the workbench the "# proteins" is sent; and this
>>> subset is re-annotated with Subsystems, for example. The "#proteins"
>>> falls
>>> since not all of the original set annotated with Genbank will be
>>> recognized
>>> here, right?
>>
>> That is right, there may be sequences in GenBank which are not present in
>> the SEED Subsystems database and vice versa.
>>
>>> But the abundance may increase because more repeats can be found in the
>>> larger database?
>>
>> When moving to a database with a larger number of sequences, the abundance
>> will be expected to increase because of the greater chance of finding a
>> match.
>>
>>> And what the "workbench abundance" means?
>> The workbench abundance is the count of the reads in the metagenomic
>> sample
>> with hits against the protein sequences (from database A, e.g. Subsystems)
>> selected in the workbench which were found in the database B (e.g.
>> GenBank).
>> In this example database A is used to select the proteins placed in the
>> workbench and database B is used to create the organism classification,
>> using the proteins from the workbench.
>>
>> Regards,
>> Mark
>>
>>
>> ----- Original Message -----
>>> From: "Gustavo B. Gregoracci" <gustavo_biomed(a)yahoo.com>
>>> To: "Mark DSouza" <dsouza(a)mcs.anl.gov>
>>> Cc: "MG- Rast" <mg-rast(a)mcs.anl.gov>, "Louisi Oliveira"
>>> <louisioliveira(a)gmail.com>
>>> Sent: Tuesday, July 19, 2011 8:46:22 PM
>>> Subject: Re: Doubt about MG-Rast annotation
>>> Hi there!
>>>
>>>
>>> I'm not sure I got it right. Let's see... I annotated the metagenome
>>> through Genbank, for example. The "# proteins" field gives me the
>>> number of unique sequences which matched the genbank database and the
>>> "abundance" is the total number of matches including repeated matches
>>> to the "# proteins" proteins. Is that so? When I send these to the
>>> workbench the "# proteins" is sent; and this subset is re-annotated
>>> with Subsystems, for example. The "#proteins" falls since not all of
>>> the original set annotated with Genbank will be recognized here,
>>> right? But the abundance may increase because more repeats can be
>>> found in the larger database?
>>>
>>>
>>>
>>> And what the "workbench abundance" means?
>>>
>>>
>>> Sorry to bother you further, but I'm could not grasp the explanation
>>> previously,
>>> Thanks again for the help and patience,
>>>
>>>
>>> Best regards,
>>> Gustavo
>>>
>>>
>>> Gustavo Bueno Gregoracci
>>> --
>>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> Email: gustavo_biomed(a)yahoo.com
>>>
>>> --
>>> Pos doc Student
>>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> Tel: +55 (21) 2562-6567
>>>
>>>
>>>
>>>
>>>
>>> From: Mark DSouza <dsouza(a)mcs.anl.gov>
>>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>>> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
>>> <louisioliveira(a)gmail.com>
>>> Sent: Monday, July 18, 2011 7:13 PM
>>> Subject: Re: Doubt about MG-Rast annotation
>>>
>>> Hi,
>>>
>>> The number discrepancy arises from the use of different databases, the
>>> original 209 hits is specific to the SEED subsystem, with 66 hits
>>> against the proteins in these organisms. When you perform the organism
>>> classification against GenBank, the number of hits is higher because
>>> of the greater number of organisms in the GenBank database, and so
>>> does the abundance counts. The workbench abundance column of the table
>>> gives a better breakdown of the taxonomic counts as compared to the
>>> 209 count. This is not very intuitive from the display and we are
>>> looking into changing the table to make this more explicit.
>>>
>>> Regards,
>>> Mark
>>>
>>> ----- Original Message -----
>>> > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>>> > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>>> > Cc: "MG- Rast" < mg-rast(a)mcs.anl.gov >, "Louisi Oliveira" <
>>> > louisioliveira(a)gmail.com >
>>> > Sent: Monday, July 18, 2011 1:03:28 PM
>>> > Subject: Re: Doubt about MG-Rast annotation
>>> > Hi there Mark,
>>> >
>>> >
>>> > Thanks for the reply. So, the metagenome ID is 4465448.3 and its
>>> > called Forno. The glitch has actually happened in others as well,
>>> > and
>>> > I'm only using this one as an example.
>>> >
>>> >
>>> > I re-checked the numbers and they still don't add. I'm working with
>>> > Genbank (e-5) for organism classification, and subsystems (e-5) for
>>> > functional classification, for that matter.
>>> >
>>> >
>>> > I have 209 hits in this entire metagenome for photosynthesis. If I
>>> > check the number of photosynthesis hits among the bacterial hits I
>>> > get
>>> > 140, while the same for eukaryotic gives me 142 hits. That puzzled
>>> > me.
>>> > When I do the opposite, and check the taxonomic composition of just
>>> > the 209 photosynthesis hits the numbers are even more weirder. I get
>>> > 245 hits for bacteria and 403 for eukaryotes!
>>> >
>>> >
>>> > Thanks again for the help,
>>> > I'm available for any further explanations necessary,
>>> >
>>> >
>>> > Best regards,
>>> > Gustavo
>>> >
>>> > Gustavo Bueno Gregoracci
>>> > --
>>> > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> > Email: gustavo_biomed(a)yahoo.com
>>> >
>>> > --
>>> > Pos doc Student
>>> > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> > Tel: +55 (21) 2562-6567
>>> >
>>> >
>>> >
>>> >
>>> >
>>> > From: Mark DSouza < dsouza(a)mcs.anl.gov >
>>> > To: Gustavo B. Gregoracci < gustavo_biomed(a)yahoo.com >
>>> > Cc: MG- Rast < mg-rast(a)mcs.anl.gov >
>>> > Sent: Monday, July 18, 2011 1:33 PM
>>> > Subject: Re: Doubt about MG-Rast annotation
>>> >
>>> > Hi,
>>> >
>>> > Can you please send us the MG-RAST ID for the dataset where you
>>> > observed this, we will check into it for you.
>>> >
>>> > --
>>> > Regards,
>>> > Mark D'Souza
>>> > -- for the MG-RAST team
>>> >
>>> > mg-rast(a)mcs.anl.gov
>>> > http://metagenomics.anl.gov/
>>> >
>>> > If you need to respond to this email please address it to the
>>> > mailing-list mg-rast(a)mcs.anl.gov
>>> >
>>> >
>>> > ----- Original Message -----
>>> > > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>>> > > To: "Mark D'Souza" < dsouza(a)mcs.anl.gov >
>>> > > Sent: Saturday, July 16, 2011 10:05:14 AM
>>> > > Subject: Doubt about MG-Rast annotation
>>> > > Hello there Mark,
>>> > >
>>> > >
>>> > >
>>> > > As I was analyzing some metagenomes, I ran into a recurrent
>>> > > problem.
>>> > > I
>>> > > was trying to cross information from taxonomic and functional
>>> > > annotation, and discovered some inconsistencies. I could not
>>> > > figure
>>> > > those out so I’m writing you in the hope that you can help me,
>>> > > since
>>> > > this will affect my analysis.
>>> > >
>>> > >
>>> > > You see… I wanted to understand the contribution of different
>>> > > phylogenetic groups to a given subsystem. So I got the total
>>> > > number
>>> > > of
>>> > > sequences for the photosynthesis subsystem, for example, which was
>>> > > 209
>>> > > hits. Then I went into taxonomy and separated all entries
>>> > > identified
>>> > > as eukaryotes, sending them to the workbench. I changed again to
>>> > > functional and checked the total number of eukaryotic hits to
>>> > > photosynthesis, which was 163. So far so good, but I decided to
>>> > > perform the same analysis regarding prokaryotes. Sent all
>>> > > prokaryotic
>>> > > entries to the workbench and changed to functional to see their
>>> > > contribution to photosynthesis as well. I got 174 hits!
>>> > >
>>> > >
>>> > >
>>> > > How is this possible? If my whole metagenome has 209 hits to a
>>> > > subsystem, how can I have 163 eukaryotic hits within it and 174
>>> > > prokaryotic hits to the same subsystem? Can sequences be annotated
>>> > > as
>>> > > both eukaryotic and prokaryotic? Cause I was interpreting these
>>> > > categories as mutually exclusive…
>>> > >
>>> > >
>>> > > Thanks in advance for the help,
>>> > > Looking forward to hear from you,
>>> > >
>>> > >
>>> > > Best regards,
>>> > > Gustavo
>>> > >
>>> > > Gustavo Bueno Gregoracci
>>> > > --
>>> > > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> > > Email: gustavo_biomed(a)yahoo.com
>>> > >
>>> > > --
>>> > > Pos doc Student
>>> > > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> > > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> > > Tel: +55 (21) 2562-6567
>>
>>
>>
>>
>>
>
>
>
1
0
Hi there
William,
Re-reading
the text it really seems a little confusing. Sorry for that. But let’s try
again:
Let me make sure we're doing exactly the same
thing.
Go to analysis page, Functional annotation tab
Select 4465448.3 with the default annotation source subsystems
Generate a table
group by level 1
select photosynthesis, with
abundance=221 and proteins=74
click To workbench.
Ok. Check. :)
What are your next steps?
Change to “organism
classification”.
Use “GenBank”
as “Annotation sources”.
Mark the “use
proteins from workbench” option.
Generate a
Domain Level table.
Now we have
only two domains, which will make it easier to explain. Mark had explained us
the “abundance” and “hits” fields like you did. We understood it when they
refer to a metagenome, but they make less sense when working with the
Workbench.
We imagined
that when you select things to send to the workbench and check the “use
proteins from workbench” option, the analysis would be restricted to the set of
things sent to the workbench, in this case to the 74 unique proteins and 221 total
hits. So we get 56 bacterial proteins and 101 eukaryotic ones,
adding up to 157, instead of the total 221 sent. But this loss is
understandable since the databases are different and possibly not all 221 are
identified in the other database. First thing is to know if this interpretation
is right.
Second doubt
is about the “#hits” in this domain table, or the total number of unique sequences generated from the workbench. If our above explanation is correct,
we suppose the analysis would be restricted to the 221 hits. How do they turn
into the new “#hits” of this new table? Let’s get to the numbers… From the 221
hits sent to the workbench, 157 were identified (“workbench abundance” bacteria
+ eukarya). How this becomes 97 unique proteins for bacteria and 163 for
eukarya (“#hits”)? 163+97 add up to the 260 which is more than the total of
hits sent to the workbench (221) and the total identified (157).
Similarly,
how the new “abundance” is generated? If we sent 221 sequences to the
workbench, and 157 were identified, where do 245 bacterial and 403 eukaryotic
sequences come from? They add up to 648!
Sorry to bother you,
Thanks for
the help!!
Best
regards!
Gustavo
Gustavo Bueno Gregoracci
--
MSc., Dr. in Microbiology (Genetics and Molecular Biology)
Email: gustavo_biomed(a)yahoo.com
--
Pos doc Student
Laboratório de Microbiologia (Prof. Fabiano Thompson)
Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
Tel: +55 (21) 2562-6567
>________________________________
>From: William Trimble <wltrimbl(a)gmail.com>
>To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>; mg-rast(a)mcs.anl.gov
>Sent: Friday, July 29, 2011 2:32 PM
>Subject: Re: [mg-rast] Doubt about MG-Rast annotation
>
>Sorry for the delay in getting back to you.
>
>> Let's use an example, like the metagenome I sent you. We used the Subsytems
>> database and separated Photosynthesis with 209 of abundance and 66 unique
>> proteins in the group. We sent it to the Workbench to look into the
>> taxonomic profile (Genbank).
>
>This means that there were 66 distinct sequences in the database that were
>hit by a total of 209 predicted genes in your sequence data.
>
>> 1. It stored all the 209 hits from the abundance in the workbench, right?
>> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
>> Bacteria, which is shown in the workbench abundance column. Is this
>> reasoning right? This loss is the result of the conversion between
>> databases?
>> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
>> each sequence generate more than one annotation? Cause we got for bacteria
>> an abundance of 245 and for eukarya an abundance of 403, from the 209
>> sequences selected in the workbench.
>> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
>> from the 209 hits (see question 1), how do they turn into 97 unique proteins
>> for bacteria and 163 unique proteins for eukarya?
>This is a bit of a puzzle.
>
>Let me make sure we're doing exactly the same thing.
>Go to analysis page, Functional annotation tab
>Select 4465448.3 with the default annotation source subsystems
>Generate a table
>group by level 1
>select photosynthesis, with
>abundance=221 and proteins=74
>click To workbench.
>What are your next steps?
>
>The "abundance" field is intended to be proportional to the number of times a
>given annotation is present in the data; the "hits" only counts unique
>sequences in the database.
>
>I'm not immediately sure why subsets of the sequences on the workbench
>have unexpected sizes.
>
>William Trimble
>--for the MG-RAST team
>
>> Again, I thank you for the time and patience,
>> Looking forward to hear from you,
>> Best regards,
>> Gustavo
>>
>> Gustavo Bueno Gregoracci
>> --
>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> Email: gustavo_biomed(a)yahoo.com
>> --
>> Pos doc Student
>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> Tel: +55 (21) 2562-6567
>>
>> ________________________________
>> From: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>> To: Mark DSouza <dsouza(a)mcs.anl.gov>
>> Cc: Louisi Oliveira <louisioliveira(a)gmail.com>; MG- Rast
>> <mg-rast(a)mcs.anl.gov>
>> Sent: Wednesday, July 20, 2011 5:53 PM
>> Subject: Re: Doubt about MG-Rast annotation
>>
>> Hi there,
>> Thanks for the help so far. We cleared some important questions but there
>> are new ones.
>>
>> Let's use an example, like the metagenome I sent you. We used the Subsytems
>> database and separated Photosynthesis with 209 of abundance and 66 unique
>> proteins in the group. We sent it to the Workbench to look into the
>> taxonomic profile (Genbank).
>> 1. It stored all the 209 hits from the abundance in the workbench, right?
>> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
>> Bacteria, which is shown in the workbench abundance column. Is this
>> reasoning right? This is the conversion between databases?
>>
>> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
>> each sequence generate more than one annotation? Cause we got for bacteria
>> an abundance of 245 and for eukarya an abundance of 403, from the 209
>> sequences selected in the workbench.
>> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
>> from the 209 hits (see question 1), how do they turn into 97 unique proteins
>> for bacteria and 163 unique proteins for eukarya?
>>
>> Sorry to trouble you again, but we really need to clarify these questions to
>> continue with our analyses. We also will be spreading this information among
>> our colleagues.
>>
>> Thanks again!
>> Best regards!
>> Gustavo
>>
>> Gustavo Bueno Gregoracci
>> --
>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> Email: gustavo_biomed(a)yahoo.com
>> --
>> Pos doc Student
>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> Tel: +55 (21) 2562-6567
>>
>> ________________________________
>> From: Mark DSouza <dsouza(a)mcs.anl.gov>
>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
>> <louisioliveira(a)gmail.com>
>> Sent: Wednesday, July 20, 2011 2:44 PM
>> Subject: Re: Doubt about MG-Rast annotation
>>
>> Hi,
>>
>>> The "# proteins" field gives me the number of unique sequences which
>>> matched the genbank database ...
>>
>> No, this is not accurate, the '# proteins' field gives the number of unique
>> sequences within the selected grouping so if you group by phylum you will
>> get unique sequences within the phylum.
>>
>>> ... the "abundance" is the total number of matches including repeated
>>> matches to the "# proteins" proteins. Is that so?
>>
>> Yes, this is correct.
>>
>>> When I send these to the workbench the "# proteins" is sent; and this
>>> subset is re-annotated with Subsystems, for example. The "#proteins" falls
>>> since not all of the original set annotated with Genbank will be recognized
>>> here, right?
>>
>> That is right, there may be sequences in GenBank which are not present in
>> the SEED Subsystems database and vice versa.
>>
>>> But the abundance may increase because more repeats can be found in the
>>> larger database?
>>
>> When moving to a database with a larger number of sequences, the abundance
>> will be expected to increase because of the greater chance of finding a
>> match.
>>
>>> And what the "workbench abundance" means?
>> The workbench abundance is the count of the reads in the metagenomic sample
>> with hits against the protein sequences (from database A, e.g. Subsystems)
>> selected in the workbench which were found in the database B (e.g. GenBank).
>> In this example database A is used to select the proteins placed in the
>> workbench and database B is used to create the organism classification,
>> using the proteins from the workbench.
>>
>> Regards,
>> Mark
>>
>>
>> ----- Original Message -----
>>> From: "Gustavo B. Gregoracci" <gustavo_biomed(a)yahoo.com>
>>> To: "Mark DSouza" <dsouza(a)mcs.anl.gov>
>>> Cc: "MG- Rast" <mg-rast(a)mcs.anl.gov>, "Louisi Oliveira"
>>> <louisioliveira(a)gmail.com>
>>> Sent: Tuesday, July 19, 2011 8:46:22 PM
>>> Subject: Re: Doubt about MG-Rast annotation
>>> Hi there!
>>>
>>>
>>> I'm not sure I got it right. Let's see... I annotated the metagenome
>>> through Genbank, for example. The "# proteins" field gives me the
>>> number of unique sequences which matched the genbank database and the
>>> "abundance" is the total number of matches including repeated matches
>>> to the "# proteins" proteins. Is that so? When I send these to the
>>> workbench the "# proteins" is sent; and this subset is re-annotated
>>> with Subsystems, for example. The "#proteins" falls since not all of
>>> the original set annotated with Genbank will be recognized here,
>>> right? But the abundance may increase because more repeats can be
>>> found in the larger database?
>>>
>>>
>>>
>>> And what the "workbench abundance" means?
>>>
>>>
>>> Sorry to bother you further, but I'm could not grasp the explanation
>>> previously,
>>> Thanks again for the help and patience,
>>>
>>>
>>> Best regards,
>>> Gustavo
>>>
>>>
>>> Gustavo Bueno Gregoracci
>>> --
>>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> Email: gustavo_biomed(a)yahoo.com
>>>
>>> --
>>> Pos doc Student
>>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> Tel: +55 (21) 2562-6567
>>>
>>>
>>>
>>>
>>>
>>> From: Mark DSouza <dsouza(a)mcs.anl.gov>
>>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>>> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
>>> <louisioliveira(a)gmail.com>
>>> Sent: Monday, July 18, 2011 7:13 PM
>>> Subject: Re: Doubt about MG-Rast annotation
>>>
>>> Hi,
>>>
>>> The number discrepancy arises from the use of different databases, the
>>> original 209 hits is specific to the SEED subsystem, with 66 hits
>>> against the proteins in these organisms. When you perform the organism
>>> classification against GenBank, the number of hits is higher because
>>> of the greater number of organisms in the GenBank database, and so
>>> does the abundance counts. The workbench abundance column of the table
>>> gives a better breakdown of the taxonomic counts as compared to the
>>> 209 count. This is not very intuitive from the display and we are
>>> looking into changing the table to make this more explicit.
>>>
>>> Regards,
>>> Mark
>>>
>>> ----- Original Message -----
>>> > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>>> > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>>> > Cc: "MG- Rast" < mg-rast(a)mcs.anl.gov >, "Louisi Oliveira" <
>>> > louisioliveira(a)gmail.com >
>>> > Sent: Monday, July 18, 2011 1:03:28 PM
>>> > Subject: Re: Doubt about MG-Rast annotation
>>> > Hi there Mark,
>>> >
>>> >
>>> > Thanks for the reply. So, the metagenome ID is 4465448.3 and its
>>> > called Forno. The glitch has actually happened in others as well,
>>> > and
>>> > I'm only using this one as an example.
>>> >
>>> >
>>> > I re-checked the numbers and they still don't add. I'm working with
>>> > Genbank (e-5) for organism classification, and subsystems (e-5) for
>>> > functional classification, for that matter.
>>> >
>>> >
>>> > I have 209 hits in this entire metagenome for photosynthesis. If I
>>> > check the number of photosynthesis hits among the bacterial hits I
>>> > get
>>> > 140, while the same for eukaryotic gives me 142 hits. That puzzled
>>> > me.
>>> > When I do the opposite, and check the taxonomic composition of just
>>> > the 209 photosynthesis hits the numbers are even more weirder. I get
>>> > 245 hits for bacteria and 403 for eukaryotes!
>>> >
>>> >
>>> > Thanks again for the help,
>>> > I'm available for any further explanations necessary,
>>> >
>>> >
>>> > Best regards,
>>> > Gustavo
>>> >
>>> > Gustavo Bueno Gregoracci
>>> > --
>>> > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> > Email: gustavo_biomed(a)yahoo.com
>>> >
>>> > --
>>> > Pos doc Student
>>> > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> > Tel: +55 (21) 2562-6567
>>> >
>>> >
>>> >
>>> >
>>> >
>>> > From: Mark DSouza < dsouza(a)mcs.anl.gov >
>>> > To: Gustavo B. Gregoracci < gustavo_biomed(a)yahoo.com >
>>> > Cc: MG- Rast < mg-rast(a)mcs.anl.gov >
>>> > Sent: Monday, July 18, 2011 1:33 PM
>>> > Subject: Re: Doubt about MG-Rast annotation
>>> >
>>> > Hi,
>>> >
>>> > Can you please send us the MG-RAST ID for the dataset where you
>>> > observed this, we will check into it for you.
>>> >
>>> > --
>>> > Regards,
>>> > Mark D'Souza
>>> > -- for the MG-RAST team
>>> >
>>> > mg-rast(a)mcs.anl.gov
>>> > http://metagenomics.anl.gov/
>>> >
>>> > If you need to respond to this email please address it to the
>>> > mailing-list mg-rast(a)mcs.anl.gov
>>> >
>>> >
>>> > ----- Original Message -----
>>> > > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>>> > > To: "Mark D'Souza" < dsouza(a)mcs.anl.gov >
>>> > > Sent: Saturday, July 16, 2011 10:05:14 AM
>>> > > Subject: Doubt about MG-Rast annotation
>>> > > Hello there Mark,
>>> > >
>>> > >
>>> > >
>>> > > As I was analyzing some metagenomes, I ran into a recurrent
>>> > > problem.
>>> > > I
>>> > > was trying to cross information from taxonomic and functional
>>> > > annotation, and discovered some inconsistencies. I could not
>>> > > figure
>>> > > those out so I’m writing you in the hope that you can help me,
>>> > > since
>>> > > this will affect my analysis.
>>> > >
>>> > >
>>> > > You see… I wanted to understand the contribution of different
>>> > > phylogenetic groups to a given subsystem. So I got the total
>>> > > number
>>> > > of
>>> > > sequences for the photosynthesis subsystem, for example, which was
>>> > > 209
>>> > > hits. Then I went into taxonomy and separated all entries
>>> > > identified
>>> > > as eukaryotes, sending them to the workbench. I changed again to
>>> > > functional and checked the total number of eukaryotic hits to
>>> > > photosynthesis, which was 163. So far so good, but I decided to
>>> > > perform the same analysis regarding prokaryotes. Sent all
>>> > > prokaryotic
>>> > > entries to the workbench and changed to functional to see their
>>> > > contribution to photosynthesis as well. I got 174 hits!
>>> > >
>>> > >
>>> > >
>>> > > How is this possible? If my whole metagenome has 209 hits to a
>>> > > subsystem, how can I have 163 eukaryotic hits within it and 174
>>> > > prokaryotic hits to the same subsystem? Can sequences be annotated
>>> > > as
>>> > > both eukaryotic and prokaryotic? Cause I was interpreting these
>>> > > categories as mutually exclusive…
>>> > >
>>> > >
>>> > > Thanks in advance for the help,
>>> > > Looking forward to hear from you,
>>> > >
>>> > >
>>> > > Best regards,
>>> > > Gustavo
>>> > >
>>> > > Gustavo Bueno Gregoracci
>>> > > --
>>> > > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>>> > > Email: gustavo_biomed(a)yahoo.com
>>> > >
>>> > > --
>>> > > Pos doc Student
>>> > > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>>> > > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>>> > > Tel: +55 (21) 2562-6567
>>
>>
>>
>>
>>
>
>
>
1
0
Your job looks like it will finish without problems.
Let us know if it isn't done by the middle of next week.
William Trimble
--for the MG-RAST team
On Fri, Jul 29, 2011 at 10:23 AM, Omry Koren <korenomry(a)gmail.com> wrote:
> Hi Mark,
> Thanks for all the previous help. At the moment I have one file stuck at
> finalizing for a long time, is there something I can do?
> Thanks
> Omry
>
> On Tue, Jul 5, 2011 at 9:57 AM, Omry Koren <korenomry(a)gmail.com> wrote:
>>
>> Hi Mark,
>> A few of my files are stuck with an error. Is there anything I should do?
>> Thank you very much for all your help and time,
>> Omry
>>
>> On Wed, Jun 29, 2011 at 12:50 PM, Mark DSouza <dsouza(a)mcs.anl.gov> wrote:
>>>
>>> Hi Omry,
>>>
>>> There is no way to delete jobs from the web interface right now, we will
>>> be adding this in the future. In the meantime I can delete the jobs, is this
>>> the correct list?
>>> 25749 4466152.3 1234
>>> 25892 4466295.3 1264
>>> 25853 4466256.3 2264
>>> 25747 4466150.3 2264
>>> 25687 4466090.3 2195
>>>
>>> We will check the job 1303 (25745, 4466148.3) which is showing an error.
>>>
>>> Regards,
>>> Mark
>>>
>>>
>>> ----- Original Message -----
>>> > From: "Omry Koren" <korenomry(a)gmail.com>
>>> > To: mg-rast(a)mcs.anl.gov
>>> > Sent: Wednesday, June 29, 2011 10:36:21 AM
>>> > Subject: [mg-rast] Fwd: Submission to MG-RAST
>>> > ---------- Forwarded message ----------
>>> > From: Omry Koren < korenomry(a)gmail.com >
>>> > Date: Tue, Jun 21, 2011 at 12:18 PM
>>> > Subject: Re: Submission to MG-RAST
>>> > To: Mark DSouza < dsouza(a)mcs.anl.gov >
>>> >
>>> > Hi Mark,
>>> > I went over all of the files and I have a few questions:
>>> > 1. How do I delete files? I have a few duplicates (1234,1264, and
>>> > 2195) and I would also like to delete both 2264 since the sequencing
>>> > center just notified me that they are very bad quality.
>>> > 2. A few files are showing errors, anything I can do?
>>> >
>>> >
>>> >
>>> >
>>> > Thank you very much for your time,
>>> > Omry
>>> >
>>> >
>>> >
>>> >
>>> >
>>> > On Mon, Jun 13, 2011 at 8:24 PM, Omry Koren < korenomry(a)gmail.com >
>>> > wrote:
>>> >
>>> >
>>> > I am really sorry about that, I thought some of the files were
>>> > corrupted. This is my first time using MG-RAST and I will contact you
>>> > next time I run into trouble.
>>> > Thank you and sorry again,
>>> > Omry
>>> >
>>> >
>>> >
>>> >
>>> >
>>> > On Mon, Jun 13, 2011 at 5:32 PM, Mark DSouza < dsouza(a)mcs.anl.gov >
>>> > wrote:
>>> >
>>> >
>>> > Hi,
>>> >
>>> > I am still seeing multiple uploads for the same file, 2195 was
>>> > uploaded four times, including once today, 2264 has been uploaded
>>> > thrice, including once today, 1234 and 1264 were uploaded twice with
>>> > an upload of 1264 today. Check if a dataset is already running before
>>> > uploading it and contact us if a job is not progressing, we can fix
>>> > things on our side. Creating duplicate jobs loads our compute machines
>>> > unnecessarily and affects the processing times for all our users.
>>> > These datasets are each multiple Gbp in size and will not get
>>> > completed in a day, if you are patient and work with us the jobs will
>>> > get processed.
>>> >
>>> >
>>> > Regards,
>>> > Mark
>>> >
>>> >
>>> >
>>> > ----- Original Message -----
>>> > > From: "Omry Koren" < korenomry(a)gmail.com >
>>> > > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>>> >
>>> >
>>> >
>>> > > Sent: Friday, June 10, 2011 5:09:17 PM
>>> > > Subject: Re: Submission to MG-RAST
>>> > > okay, thank you very much
>>> > >
>>> > >
>>> > > On Fri, Jun 10, 2011 at 6:06 PM, Mark DSouza < dsouza(a)mcs.anl.gov >
>>> > > wrote:
>>> > >
>>> > >
>>> > > Hi,
>>> > >
>>> > > No problem. I deleted the duplicated jobs. I also noticed that the
>>> > > upload for 2342 seems to be truncated, please compare the file size
>>> > > from the upload page against your local file, you probably need to
>>> > > redo the upload for this job.
>>> > >
>>> > > Regards,
>>> > > Mark
>>> > >
>>> > >
>>> > >
>>> > >
>>> > >
>>> > > ----- Original Message -----
>>> > > > From: "Omry Koren" < korenomry(a)gmail.com >
>>> > > > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>>> > > > Sent: Friday, June 10, 2011 3:26:17 PM
>>> > > > Subject: Re: Submission to MG-RAST
>>> > > > Hi,
>>> > > > Yes those are duplicates and can be deleted, thank you.
>>> > > > Besides what I am uploading now I only have one more file to
>>> > > > upload
>>> > > > so
>>> > > > will save the ftp option for next time.
>>> > > > Thank you very much
>>> > > > Omry
>>> > > >
>>> > > >
>>> > > > On Fri, Jun 10, 2011 at 4:24 PM, Mark DSouza < dsouza(a)mcs.anl.gov
>>> > > > >
>>> > > > wrote:
>>> > > >
>>> > > >
>>> > > > Hi,
>>> > > >
>>> > > > It appears that you are uploading a large amount of sequence data
>>> > > > to
>>> > > > MG-RAST. While this is not a problem, it would make the process
>>> > > > easier
>>> > > > for both of us if we coordinate these uploads. If you have more
>>> > > > data
>>> > > > to submit we can create an ftp site where you can dump your data
>>> > > > files
>>> > > > and we can then use the files to complete the submission.
>>> > > >
>>> > > > From the jobs which have been created the files 2195.txt and
>>> > > > 2264.txt
>>> > > > have been uploaded multiple times and seem to be running multiple
>>> > > > jobs. If you can confirm that these are identical datasets I will
>>> > > > delete the duplicated jobs.
>>> > > >
>>> > > > Please let us know about the ftp option, this will make the upload
>>> > > > process easier for you as well as help us in balancing the load on
>>> > > > our
>>> > > > machines and avoid duplicating jobs.
>>> > > >
>>> > > > --
>>> > > > Regards,
>>> > > > Mark D'Souza
>>> > > > -- for the MG-RAST team
>>> > > >
>>> > > > mg-rast(a)mcs.anl.gov
>>> > > > http://metagenomics.anl.gov/
>>> > > >
>>> > > > If you need to respond to this email please address it to the
>>> > > > mailing-list mg-rast(a)mcs.anl.gov
>>> > > >
>>> > > >
>>> > > >
>>> > > > --
>>> > > > Omry Koren, Ph.D.
>>> > > > Postdoctoral Associate
>>> > > > Department of Microbiology
>>> > > > 467 Biotechnology building
>>> > > > Cornell University
>>> > > > Ithaca, NY 14853
>>> > > > Email: ok46(a)cornell.edu
>>> > >
>>> > >
>>> > >
>>> > > --
>>> > > Omry Koren, Ph.D.
>>> > > Postdoctoral Associate
>>> > > Department of Microbiology
>>> > > 467 Biotechnology building
>>> > > Cornell University
>>> > > Ithaca, NY 14853
>>> > > Email: ok46(a)cornell.edu
>>> >
>>> >
>>> >
>>> > --
>>> >
>>> > Omry Koren, Ph.D.
>>> > Postdoctoral Associate
>>> > Department of Microbiology
>>> > 467 Biotechnology building
>>> > Cornell University
>>> > Ithaca, NY 14853
>>> > Email: ok46(a)cornell.edu
>>> >
>>> >
>>> >
>>> >
>>> > --
>>> >
>>> > Omry Koren, Ph.D.
>>> > Postdoctoral Associate
>>> > Department of Microbiology
>>> > 467 Biotechnology building
>>> > Cornell University
>>> > Ithaca, NY 14853
>>> > Email: ok46(a)cornell.edu
>>> >
>>> >
>>> >
>>> >
>>> > --
>>> > Omry Koren, Ph.D.
>>> > Postdoctoral Associate
>>> > Department of Microbiology
>>> > 467 Biotechnology building
>>> > Cornell University
>>> > Ithaca, NY 14853
>>> > Email: ok46(a)cornell.edu
>>
>>
>>
>> --
>> Omry Koren, Ph.D.
>> Postdoctoral Associate
>> Department of Microbiology
>> 467 Biotechnology building
>> Cornell University
>> Ithaca, NY 14853
>> Email: ok46(a)cornell.edu
>>
>
>
>
> --
> Omry Koren, Ph.D.
> Postdoctoral Associate
> Department of Microbiology
> 467 Biotechnology building
> Cornell University
> Ithaca, NY 14853
> Email: ok46(a)cornell.edu
>
>
1
0
Its fine, sims are being parsed right now.
-travis
On Fri, Jul 29, 2011 at 11:09 AM, William Trimble <wltrimbl(a)gmail.com>wrote:
> This is job 26706
>
> Does this require action?
>
> w
>
> ---------- Forwarded message ----------
> From: Omry Koren <korenomry(a)gmail.com>
> Date: Fri, Jul 29, 2011 at 10:23 AM
> Subject: Re: [mg-rast] Fwd: Submission to MG-RAST
> To: Mark DSouza <dsouza(a)mcs.anl.gov>
> Cc: mg-rast(a)mcs.anl.gov
>
>
> Hi Mark,
> Thanks for all the previous help. At the moment I have one file stuck
> at finalizing for a long time, is there something I can do?
> Thanks
> Omry
>
> On Tue, Jul 5, 2011 at 9:57 AM, Omry Koren <korenomry(a)gmail.com> wrote:
> >
> > Hi Mark,
> > A few of my files are stuck with an error. Is there anything I should do?
> > Thank you very much for all your help and time,
> > Omry
> >
> > On Wed, Jun 29, 2011 at 12:50 PM, Mark DSouza <dsouza(a)mcs.anl.gov>
> wrote:
> >>
> >> Hi Omry,
> >>
> >> There is no way to delete jobs from the web interface right now, we will
> be adding this in the future. In the meantime I can delete the jobs, is this
> the correct list?
> >> 25749 4466152.3 1234
> >> 25892 4466295.3 1264
> >> 25853 4466256.3 2264
> >> 25747 4466150.3 2264
> >> 25687 4466090.3 2195
> >>
> >> We will check the job 1303 (25745, 4466148.3) which is showing an error.
> >>
> >> Regards,
> >> Mark
> >>
> >>
> >> ----- Original Message -----
> >> > From: "Omry Koren" <korenomry(a)gmail.com>
> >> > To: mg-rast(a)mcs.anl.gov
> >> > Sent: Wednesday, June 29, 2011 10:36:21 AM
> >> > Subject: [mg-rast] Fwd: Submission to MG-RAST
> >> > ---------- Forwarded message ----------
> >> > From: Omry Koren < korenomry(a)gmail.com >
> >> > Date: Tue, Jun 21, 2011 at 12:18 PM
> >> > Subject: Re: Submission to MG-RAST
> >> > To: Mark DSouza < dsouza(a)mcs.anl.gov >
> >> >
> >> > Hi Mark,
> >> > I went over all of the files and I have a few questions:
> >> > 1. How do I delete files? I have a few duplicates (1234,1264, and
> >> > 2195) and I would also like to delete both 2264 since the sequencing
> >> > center just notified me that they are very bad quality.
> >> > 2. A few files are showing errors, anything I can do?
> >> >
> >> >
> >> >
> >> >
> >> > Thank you very much for your time,
> >> > Omry
> >> >
> >> >
> >> >
> >> >
> >> >
> >> > On Mon, Jun 13, 2011 at 8:24 PM, Omry Koren < korenomry(a)gmail.com >
> >> > wrote:
> >> >
> >> >
> >> > I am really sorry about that, I thought some of the files were
> >> > corrupted. This is my first time using MG-RAST and I will contact you
> >> > next time I run into trouble.
> >> > Thank you and sorry again,
> >> > Omry
> >> >
> >> >
> >> >
> >> >
> >> >
> >> > On Mon, Jun 13, 2011 at 5:32 PM, Mark DSouza < dsouza(a)mcs.anl.gov >
> >> > wrote:
> >> >
> >> >
> >> > Hi,
> >> >
> >> > I am still seeing multiple uploads for the same file, 2195 was
> >> > uploaded four times, including once today, 2264 has been uploaded
> >> > thrice, including once today, 1234 and 1264 were uploaded twice with
> >> > an upload of 1264 today. Check if a dataset is already running before
> >> > uploading it and contact us if a job is not progressing, we can fix
> >> > things on our side. Creating duplicate jobs loads our compute machines
> >> > unnecessarily and affects the processing times for all our users.
> >> > These datasets are each multiple Gbp in size and will not get
> >> > completed in a day, if you are patient and work with us the jobs will
> >> > get processed.
> >> >
> >> >
> >> > Regards,
> >> > Mark
> >> >
> >> >
> >> >
> >> > ----- Original Message -----
> >> > > From: "Omry Koren" < korenomry(a)gmail.com >
> >> > > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
> >> >
> >> >
> >> >
> >> > > Sent: Friday, June 10, 2011 5:09:17 PM
> >> > > Subject: Re: Submission to MG-RAST
> >> > > okay, thank you very much
> >> > >
> >> > >
> >> > > On Fri, Jun 10, 2011 at 6:06 PM, Mark DSouza < dsouza(a)mcs.anl.gov >
> >> > > wrote:
> >> > >
> >> > >
> >> > > Hi,
> >> > >
> >> > > No problem. I deleted the duplicated jobs. I also noticed that the
> >> > > upload for 2342 seems to be truncated, please compare the file size
> >> > > from the upload page against your local file, you probably need to
> >> > > redo the upload for this job.
> >> > >
> >> > > Regards,
> >> > > Mark
> >> > >
> >> > >
> >> > >
> >> > >
> >> > >
> >> > > ----- Original Message -----
> >> > > > From: "Omry Koren" < korenomry(a)gmail.com >
> >> > > > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
> >> > > > Sent: Friday, June 10, 2011 3:26:17 PM
> >> > > > Subject: Re: Submission to MG-RAST
> >> > > > Hi,
> >> > > > Yes those are duplicates and can be deleted, thank you.
> >> > > > Besides what I am uploading now I only have one more file to
> >> > > > upload
> >> > > > so
> >> > > > will save the ftp option for next time.
> >> > > > Thank you very much
> >> > > > Omry
> >> > > >
> >> > > >
> >> > > > On Fri, Jun 10, 2011 at 4:24 PM, Mark DSouza < dsouza(a)mcs.anl.gov
> >> > > > >
> >> > > > wrote:
> >> > > >
> >> > > >
> >> > > > Hi,
> >> > > >
> >> > > > It appears that you are uploading a large amount of sequence data
> >> > > > to
> >> > > > MG-RAST. While this is not a problem, it would make the process
> >> > > > easier
> >> > > > for both of us if we coordinate these uploads. If you have more
> >> > > > data
> >> > > > to submit we can create an ftp site where you can dump your data
> >> > > > files
> >> > > > and we can then use the files to complete the submission.
> >> > > >
> >> > > > From the jobs which have been created the files 2195.txt and
> >> > > > 2264.txt
> >> > > > have been uploaded multiple times and seem to be running multiple
> >> > > > jobs. If you can confirm that these are identical datasets I will
> >> > > > delete the duplicated jobs.
> >> > > >
> >> > > > Please let us know about the ftp option, this will make the upload
> >> > > > process easier for you as well as help us in balancing the load on
> >> > > > our
> >> > > > machines and avoid duplicating jobs.
> >> > > >
> >> > > > --
> >> > > > Regards,
> >> > > > Mark D'Souza
> >> > > > -- for the MG-RAST team
> >> > > >
> >> > > > mg-rast(a)mcs.anl.gov
> >> > > > http://metagenomics.anl.gov/
> >> > > >
> >> > > > If you need to respond to this email please address it to the
> >> > > > mailing-list mg-rast(a)mcs.anl.gov
> >> > > >
> >> > > >
> >> > > >
> >> > > > --
> >> > > > Omry Koren, Ph.D.
> >> > > > Postdoctoral Associate
> >> > > > Department of Microbiology
> >> > > > 467 Biotechnology building
> >> > > > Cornell University
> >> > > > Ithaca, NY 14853
> >> > > > Email: ok46(a)cornell.edu
> >> > >
> >> > >
> >> > >
> >> > > --
> >> > > Omry Koren, Ph.D.
> >> > > Postdoctoral Associate
> >> > > Department of Microbiology
> >> > > 467 Biotechnology building
> >> > > Cornell University
> >> > > Ithaca, NY 14853
> >> > > Email: ok46(a)cornell.edu
> >> >
> >> >
> >> >
> >> > --
> >> >
> >> > Omry Koren, Ph.D.
> >> > Postdoctoral Associate
> >> > Department of Microbiology
> >> > 467 Biotechnology building
> >> > Cornell University
> >> > Ithaca, NY 14853
> >> > Email: ok46(a)cornell.edu
> >> >
> >> >
> >> >
> >> >
> >> > --
> >> >
> >> > Omry Koren, Ph.D.
> >> > Postdoctoral Associate
> >> > Department of Microbiology
> >> > 467 Biotechnology building
> >> > Cornell University
> >> > Ithaca, NY 14853
> >> > Email: ok46(a)cornell.edu
> >> >
> >> >
> >> >
> >> >
> >> > --
> >> > Omry Koren, Ph.D.
> >> > Postdoctoral Associate
> >> > Department of Microbiology
> >> > 467 Biotechnology building
> >> > Cornell University
> >> > Ithaca, NY 14853
> >> > Email: ok46(a)cornell.edu
> >
> >
> >
> > --
> > Omry Koren, Ph.D.
> > Postdoctoral Associate
> > Department of Microbiology
> > 467 Biotechnology building
> > Cornell University
> > Ithaca, NY 14853
> > Email: ok46(a)cornell.edu
> >
>
>
>
> --
> Omry Koren, Ph.D.
> Postdoctoral Associate
> Department of Microbiology
> 467 Biotechnology building
> Cornell University
> Ithaca, NY 14853
> Email: ok46(a)cornell.edu
>
1
0
Sorry for the delay in getting back to you.
> Let's use an example, like the metagenome I sent you. We used the Subsytems
> database and separated Photosynthesis with 209 of abundance and 66 unique
> proteins in the group. We sent it to the Workbench to look into the
> taxonomic profile (Genbank).
This means that there were 66 distinct sequences in the database that were
hit by a total of 209 predicted genes in your sequence data.
> 1. It stored all the 209 hits from the abundance in the workbench, right?
> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
> Bacteria, which is shown in the workbench abundance column. Is this
> reasoning right? This loss is the result of the conversion between
> databases?
> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
> each sequence generate more than one annotation? Cause we got for bacteria
> an abundance of 245 and for eukarya an abundance of 403, from the 209
> sequences selected in the workbench.
> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
> from the 209 hits (see question 1), how do they turn into 97 unique proteins
> for bacteria and 163 unique proteins for eukarya?
This is a bit of a puzzle.
Let me make sure we're doing exactly the same thing.
Go to analysis page, Functional annotation tab
Select 4465448.3 with the default annotation source subsystems
Generate a table
group by level 1
select photosynthesis, with
abundance=221 and proteins=74
click To workbench.
What are your next steps?
The "abundance" field is intended to be proportional to the number of times a
given annotation is present in the data; the "hits" only counts unique
sequences in the database.
I'm not immediately sure why subsets of the sequences on the workbench
have unexpected sizes.
William Trimble
--for the MG-RAST team
> Again, I thank you for the time and patience,
> Looking forward to hear from you,
> Best regards,
> Gustavo
>
> Gustavo Bueno Gregoracci
> --
> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
> Email: gustavo_biomed(a)yahoo.com
> --
> Pos doc Student
> Laboratório de Microbiologia (Prof. Fabiano Thompson)
> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
> Tel: +55 (21) 2562-6567
>
> ________________________________
> From: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
> To: Mark DSouza <dsouza(a)mcs.anl.gov>
> Cc: Louisi Oliveira <louisioliveira(a)gmail.com>; MG- Rast
> <mg-rast(a)mcs.anl.gov>
> Sent: Wednesday, July 20, 2011 5:53 PM
> Subject: Re: Doubt about MG-Rast annotation
>
> Hi there,
> Thanks for the help so far. We cleared some important questions but there
> are new ones.
>
> Let's use an example, like the metagenome I sent you. We used the Subsytems
> database and separated Photosynthesis with 209 of abundance and 66 unique
> proteins in the group. We sent it to the Workbench to look into the
> taxonomic profile (Genbank).
> 1. It stored all the 209 hits from the abundance in the workbench, right?
> But in the Genbank it identified only 150, 101 as Eukaryota and 49 as
> Bacteria, which is shown in the workbench abundance column. Is this
> reasoning right? This is the conversion between databases?
>
> 2. Each one of these 209 sequences is re-annotated in the Genbank, but can
> each sequence generate more than one annotation? Cause we got for bacteria
> an abundance of 245 and for eukarya an abundance of 403, from the 209
> sequences selected in the workbench.
> 3. Similarly, if we got 150 proteins recognized in the workbench abundance
> from the 209 hits (see question 1), how do they turn into 97 unique proteins
> for bacteria and 163 unique proteins for eukarya?
>
> Sorry to trouble you again, but we really need to clarify these questions to
> continue with our analyses. We also will be spreading this information among
> our colleagues.
>
> Thanks again!
> Best regards!
> Gustavo
>
> Gustavo Bueno Gregoracci
> --
> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
> Email: gustavo_biomed(a)yahoo.com
> --
> Pos doc Student
> Laboratório de Microbiologia (Prof. Fabiano Thompson)
> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
> Tel: +55 (21) 2562-6567
>
> ________________________________
> From: Mark DSouza <dsouza(a)mcs.anl.gov>
> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
> <louisioliveira(a)gmail.com>
> Sent: Wednesday, July 20, 2011 2:44 PM
> Subject: Re: Doubt about MG-Rast annotation
>
> Hi,
>
>> The "# proteins" field gives me the number of unique sequences which
>> matched the genbank database ...
>
> No, this is not accurate, the '# proteins' field gives the number of unique
> sequences within the selected grouping so if you group by phylum you will
> get unique sequences within the phylum.
>
>> ... the "abundance" is the total number of matches including repeated
>> matches to the "# proteins" proteins. Is that so?
>
> Yes, this is correct.
>
>> When I send these to the workbench the "# proteins" is sent; and this
>> subset is re-annotated with Subsystems, for example. The "#proteins" falls
>> since not all of the original set annotated with Genbank will be recognized
>> here, right?
>
> That is right, there may be sequences in GenBank which are not present in
> the SEED Subsystems database and vice versa.
>
>> But the abundance may increase because more repeats can be found in the
>> larger database?
>
> When moving to a database with a larger number of sequences, the abundance
> will be expected to increase because of the greater chance of finding a
> match.
>
>> And what the "workbench abundance" means?
> The workbench abundance is the count of the reads in the metagenomic sample
> with hits against the protein sequences (from database A, e.g. Subsystems)
> selected in the workbench which were found in the database B (e.g. GenBank).
> In this example database A is used to select the proteins placed in the
> workbench and database B is used to create the organism classification,
> using the proteins from the workbench.
>
> Regards,
> Mark
>
>
> ----- Original Message -----
>> From: "Gustavo B. Gregoracci" <gustavo_biomed(a)yahoo.com>
>> To: "Mark DSouza" <dsouza(a)mcs.anl.gov>
>> Cc: "MG- Rast" <mg-rast(a)mcs.anl.gov>, "Louisi Oliveira"
>> <louisioliveira(a)gmail.com>
>> Sent: Tuesday, July 19, 2011 8:46:22 PM
>> Subject: Re: Doubt about MG-Rast annotation
>> Hi there!
>>
>>
>> I'm not sure I got it right. Let's see... I annotated the metagenome
>> through Genbank, for example. The "# proteins" field gives me the
>> number of unique sequences which matched the genbank database and the
>> "abundance" is the total number of matches including repeated matches
>> to the "# proteins" proteins. Is that so? When I send these to the
>> workbench the "# proteins" is sent; and this subset is re-annotated
>> with Subsystems, for example. The "#proteins" falls since not all of
>> the original set annotated with Genbank will be recognized here,
>> right? But the abundance may increase because more repeats can be
>> found in the larger database?
>>
>>
>>
>> And what the "workbench abundance" means?
>>
>>
>> Sorry to bother you further, but I'm could not grasp the explanation
>> previously,
>> Thanks again for the help and patience,
>>
>>
>> Best regards,
>> Gustavo
>>
>>
>> Gustavo Bueno Gregoracci
>> --
>> MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> Email: gustavo_biomed(a)yahoo.com
>>
>> --
>> Pos doc Student
>> Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> Tel: +55 (21) 2562-6567
>>
>>
>>
>>
>>
>> From: Mark DSouza <dsouza(a)mcs.anl.gov>
>> To: Gustavo B. Gregoracci <gustavo_biomed(a)yahoo.com>
>> Cc: MG- Rast <mg-rast(a)mcs.anl.gov>; Louisi Oliveira
>> <louisioliveira(a)gmail.com>
>> Sent: Monday, July 18, 2011 7:13 PM
>> Subject: Re: Doubt about MG-Rast annotation
>>
>> Hi,
>>
>> The number discrepancy arises from the use of different databases, the
>> original 209 hits is specific to the SEED subsystem, with 66 hits
>> against the proteins in these organisms. When you perform the organism
>> classification against GenBank, the number of hits is higher because
>> of the greater number of organisms in the GenBank database, and so
>> does the abundance counts. The workbench abundance column of the table
>> gives a better breakdown of the taxonomic counts as compared to the
>> 209 count. This is not very intuitive from the display and we are
>> looking into changing the table to make this more explicit.
>>
>> Regards,
>> Mark
>>
>> ----- Original Message -----
>> > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>> > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>> > Cc: "MG- Rast" < mg-rast(a)mcs.anl.gov >, "Louisi Oliveira" <
>> > louisioliveira(a)gmail.com >
>> > Sent: Monday, July 18, 2011 1:03:28 PM
>> > Subject: Re: Doubt about MG-Rast annotation
>> > Hi there Mark,
>> >
>> >
>> > Thanks for the reply. So, the metagenome ID is 4465448.3 and its
>> > called Forno. The glitch has actually happened in others as well,
>> > and
>> > I'm only using this one as an example.
>> >
>> >
>> > I re-checked the numbers and they still don't add. I'm working with
>> > Genbank (e-5) for organism classification, and subsystems (e-5) for
>> > functional classification, for that matter.
>> >
>> >
>> > I have 209 hits in this entire metagenome for photosynthesis. If I
>> > check the number of photosynthesis hits among the bacterial hits I
>> > get
>> > 140, while the same for eukaryotic gives me 142 hits. That puzzled
>> > me.
>> > When I do the opposite, and check the taxonomic composition of just
>> > the 209 photosynthesis hits the numbers are even more weirder. I get
>> > 245 hits for bacteria and 403 for eukaryotes!
>> >
>> >
>> > Thanks again for the help,
>> > I'm available for any further explanations necessary,
>> >
>> >
>> > Best regards,
>> > Gustavo
>> >
>> > Gustavo Bueno Gregoracci
>> > --
>> > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> > Email: gustavo_biomed(a)yahoo.com
>> >
>> > --
>> > Pos doc Student
>> > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> > Tel: +55 (21) 2562-6567
>> >
>> >
>> >
>> >
>> >
>> > From: Mark DSouza < dsouza(a)mcs.anl.gov >
>> > To: Gustavo B. Gregoracci < gustavo_biomed(a)yahoo.com >
>> > Cc: MG- Rast < mg-rast(a)mcs.anl.gov >
>> > Sent: Monday, July 18, 2011 1:33 PM
>> > Subject: Re: Doubt about MG-Rast annotation
>> >
>> > Hi,
>> >
>> > Can you please send us the MG-RAST ID for the dataset where you
>> > observed this, we will check into it for you.
>> >
>> > --
>> > Regards,
>> > Mark D'Souza
>> > -- for the MG-RAST team
>> >
>> > mg-rast(a)mcs.anl.gov
>> > http://metagenomics.anl.gov/
>> >
>> > If you need to respond to this email please address it to the
>> > mailing-list mg-rast(a)mcs.anl.gov
>> >
>> >
>> > ----- Original Message -----
>> > > From: "Gustavo B. Gregoracci" < gustavo_biomed(a)yahoo.com >
>> > > To: "Mark D'Souza" < dsouza(a)mcs.anl.gov >
>> > > Sent: Saturday, July 16, 2011 10:05:14 AM
>> > > Subject: Doubt about MG-Rast annotation
>> > > Hello there Mark,
>> > >
>> > >
>> > >
>> > > As I was analyzing some metagenomes, I ran into a recurrent
>> > > problem.
>> > > I
>> > > was trying to cross information from taxonomic and functional
>> > > annotation, and discovered some inconsistencies. I could not
>> > > figure
>> > > those out so I’m writing you in the hope that you can help me,
>> > > since
>> > > this will affect my analysis.
>> > >
>> > >
>> > > You see… I wanted to understand the contribution of different
>> > > phylogenetic groups to a given subsystem. So I got the total
>> > > number
>> > > of
>> > > sequences for the photosynthesis subsystem, for example, which was
>> > > 209
>> > > hits. Then I went into taxonomy and separated all entries
>> > > identified
>> > > as eukaryotes, sending them to the workbench. I changed again to
>> > > functional and checked the total number of eukaryotic hits to
>> > > photosynthesis, which was 163. So far so good, but I decided to
>> > > perform the same analysis regarding prokaryotes. Sent all
>> > > prokaryotic
>> > > entries to the workbench and changed to functional to see their
>> > > contribution to photosynthesis as well. I got 174 hits!
>> > >
>> > >
>> > >
>> > > How is this possible? If my whole metagenome has 209 hits to a
>> > > subsystem, how can I have 163 eukaryotic hits within it and 174
>> > > prokaryotic hits to the same subsystem? Can sequences be annotated
>> > > as
>> > > both eukaryotic and prokaryotic? Cause I was interpreting these
>> > > categories as mutually exclusive…
>> > >
>> > >
>> > > Thanks in advance for the help,
>> > > Looking forward to hear from you,
>> > >
>> > >
>> > > Best regards,
>> > > Gustavo
>> > >
>> > > Gustavo Bueno Gregoracci
>> > > --
>> > > MSc., Dr. in Microbiology (Genetics and Molecular Biology)
>> > > Email: gustavo_biomed(a)yahoo.com
>> > >
>> > > --
>> > > Pos doc Student
>> > > Laboratório de Microbiologia (Prof. Fabiano Thompson)
>> > > Centro de Ciências de Saúde - Instituto de Biologia - UFRJ
>> > > Tel: +55 (21) 2562-6567
>
>
>
>
>
1
0
29 Jul '11
Dear Marcus ,
I have to check the version of the submission script you are using , we will get back to you as soon as possible.
Andreas
On Jul 29, 2011, at 11:01 AM, Marcus Claesson wrote:
> Dear List,
>
> Is it currently possible to upload fasta files of
> (barcode/primer-stripped) 16S pyrosequencing reads to my existing
> MG-RAST account using this script:
> http://qiime.sourceforge.net/scripts/submit_to_mgrast.html?highlight=submit…
>
> I have just generated an authorisation key but don't have the Project
> ID. How can I find that?
>
> Many thanks,
> Marcus
>
Andreas Wilke
Argonne National Laboratory
Mathematics and Computer Science Division
Bldg. 240 , 4F8
9700 South Cass Avenue
Argonne, IL 60439
(630) 252-3190 (phone)
(630) 252-5986.(fax)
1
0
This is job 26706
Does this require action?
w
---------- Forwarded message ----------
From: Omry Koren <korenomry(a)gmail.com>
Date: Fri, Jul 29, 2011 at 10:23 AM
Subject: Re: [mg-rast] Fwd: Submission to MG-RAST
To: Mark DSouza <dsouza(a)mcs.anl.gov>
Cc: mg-rast(a)mcs.anl.gov
Hi Mark,
Thanks for all the previous help. At the moment I have one file stuck
at finalizing for a long time, is there something I can do?
Thanks
Omry
On Tue, Jul 5, 2011 at 9:57 AM, Omry Koren <korenomry(a)gmail.com> wrote:
>
> Hi Mark,
> A few of my files are stuck with an error. Is there anything I should do?
> Thank you very much for all your help and time,
> Omry
>
> On Wed, Jun 29, 2011 at 12:50 PM, Mark DSouza <dsouza(a)mcs.anl.gov> wrote:
>>
>> Hi Omry,
>>
>> There is no way to delete jobs from the web interface right now, we will be adding this in the future. In the meantime I can delete the jobs, is this the correct list?
>> 25749 4466152.3 1234
>> 25892 4466295.3 1264
>> 25853 4466256.3 2264
>> 25747 4466150.3 2264
>> 25687 4466090.3 2195
>>
>> We will check the job 1303 (25745, 4466148.3) which is showing an error.
>>
>> Regards,
>> Mark
>>
>>
>> ----- Original Message -----
>> > From: "Omry Koren" <korenomry(a)gmail.com>
>> > To: mg-rast(a)mcs.anl.gov
>> > Sent: Wednesday, June 29, 2011 10:36:21 AM
>> > Subject: [mg-rast] Fwd: Submission to MG-RAST
>> > ---------- Forwarded message ----------
>> > From: Omry Koren < korenomry(a)gmail.com >
>> > Date: Tue, Jun 21, 2011 at 12:18 PM
>> > Subject: Re: Submission to MG-RAST
>> > To: Mark DSouza < dsouza(a)mcs.anl.gov >
>> >
>> > Hi Mark,
>> > I went over all of the files and I have a few questions:
>> > 1. How do I delete files? I have a few duplicates (1234,1264, and
>> > 2195) and I would also like to delete both 2264 since the sequencing
>> > center just notified me that they are very bad quality.
>> > 2. A few files are showing errors, anything I can do?
>> >
>> >
>> >
>> >
>> > Thank you very much for your time,
>> > Omry
>> >
>> >
>> >
>> >
>> >
>> > On Mon, Jun 13, 2011 at 8:24 PM, Omry Koren < korenomry(a)gmail.com >
>> > wrote:
>> >
>> >
>> > I am really sorry about that, I thought some of the files were
>> > corrupted. This is my first time using MG-RAST and I will contact you
>> > next time I run into trouble.
>> > Thank you and sorry again,
>> > Omry
>> >
>> >
>> >
>> >
>> >
>> > On Mon, Jun 13, 2011 at 5:32 PM, Mark DSouza < dsouza(a)mcs.anl.gov >
>> > wrote:
>> >
>> >
>> > Hi,
>> >
>> > I am still seeing multiple uploads for the same file, 2195 was
>> > uploaded four times, including once today, 2264 has been uploaded
>> > thrice, including once today, 1234 and 1264 were uploaded twice with
>> > an upload of 1264 today. Check if a dataset is already running before
>> > uploading it and contact us if a job is not progressing, we can fix
>> > things on our side. Creating duplicate jobs loads our compute machines
>> > unnecessarily and affects the processing times for all our users.
>> > These datasets are each multiple Gbp in size and will not get
>> > completed in a day, if you are patient and work with us the jobs will
>> > get processed.
>> >
>> >
>> > Regards,
>> > Mark
>> >
>> >
>> >
>> > ----- Original Message -----
>> > > From: "Omry Koren" < korenomry(a)gmail.com >
>> > > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>> >
>> >
>> >
>> > > Sent: Friday, June 10, 2011 5:09:17 PM
>> > > Subject: Re: Submission to MG-RAST
>> > > okay, thank you very much
>> > >
>> > >
>> > > On Fri, Jun 10, 2011 at 6:06 PM, Mark DSouza < dsouza(a)mcs.anl.gov >
>> > > wrote:
>> > >
>> > >
>> > > Hi,
>> > >
>> > > No problem. I deleted the duplicated jobs. I also noticed that the
>> > > upload for 2342 seems to be truncated, please compare the file size
>> > > from the upload page against your local file, you probably need to
>> > > redo the upload for this job.
>> > >
>> > > Regards,
>> > > Mark
>> > >
>> > >
>> > >
>> > >
>> > >
>> > > ----- Original Message -----
>> > > > From: "Omry Koren" < korenomry(a)gmail.com >
>> > > > To: "Mark DSouza" < dsouza(a)mcs.anl.gov >
>> > > > Sent: Friday, June 10, 2011 3:26:17 PM
>> > > > Subject: Re: Submission to MG-RAST
>> > > > Hi,
>> > > > Yes those are duplicates and can be deleted, thank you.
>> > > > Besides what I am uploading now I only have one more file to
>> > > > upload
>> > > > so
>> > > > will save the ftp option for next time.
>> > > > Thank you very much
>> > > > Omry
>> > > >
>> > > >
>> > > > On Fri, Jun 10, 2011 at 4:24 PM, Mark DSouza < dsouza(a)mcs.anl.gov
>> > > > >
>> > > > wrote:
>> > > >
>> > > >
>> > > > Hi,
>> > > >
>> > > > It appears that you are uploading a large amount of sequence data
>> > > > to
>> > > > MG-RAST. While this is not a problem, it would make the process
>> > > > easier
>> > > > for both of us if we coordinate these uploads. If you have more
>> > > > data
>> > > > to submit we can create an ftp site where you can dump your data
>> > > > files
>> > > > and we can then use the files to complete the submission.
>> > > >
>> > > > From the jobs which have been created the files 2195.txt and
>> > > > 2264.txt
>> > > > have been uploaded multiple times and seem to be running multiple
>> > > > jobs. If you can confirm that these are identical datasets I will
>> > > > delete the duplicated jobs.
>> > > >
>> > > > Please let us know about the ftp option, this will make the upload
>> > > > process easier for you as well as help us in balancing the load on
>> > > > our
>> > > > machines and avoid duplicating jobs.
>> > > >
>> > > > --
>> > > > Regards,
>> > > > Mark D'Souza
>> > > > -- for the MG-RAST team
>> > > >
>> > > > mg-rast(a)mcs.anl.gov
>> > > > http://metagenomics.anl.gov/
>> > > >
>> > > > If you need to respond to this email please address it to the
>> > > > mailing-list mg-rast(a)mcs.anl.gov
>> > > >
>> > > >
>> > > >
>> > > > --
>> > > > Omry Koren, Ph.D.
>> > > > Postdoctoral Associate
>> > > > Department of Microbiology
>> > > > 467 Biotechnology building
>> > > > Cornell University
>> > > > Ithaca, NY 14853
>> > > > Email: ok46(a)cornell.edu
>> > >
>> > >
>> > >
>> > > --
>> > > Omry Koren, Ph.D.
>> > > Postdoctoral Associate
>> > > Department of Microbiology
>> > > 467 Biotechnology building
>> > > Cornell University
>> > > Ithaca, NY 14853
>> > > Email: ok46(a)cornell.edu
>> >
>> >
>> >
>> > --
>> >
>> > Omry Koren, Ph.D.
>> > Postdoctoral Associate
>> > Department of Microbiology
>> > 467 Biotechnology building
>> > Cornell University
>> > Ithaca, NY 14853
>> > Email: ok46(a)cornell.edu
>> >
>> >
>> >
>> >
>> > --
>> >
>> > Omry Koren, Ph.D.
>> > Postdoctoral Associate
>> > Department of Microbiology
>> > 467 Biotechnology building
>> > Cornell University
>> > Ithaca, NY 14853
>> > Email: ok46(a)cornell.edu
>> >
>> >
>> >
>> >
>> > --
>> > Omry Koren, Ph.D.
>> > Postdoctoral Associate
>> > Department of Microbiology
>> > 467 Biotechnology building
>> > Cornell University
>> > Ithaca, NY 14853
>> > Email: ok46(a)cornell.edu
>
>
>
> --
> Omry Koren, Ph.D.
> Postdoctoral Associate
> Department of Microbiology
> 467 Biotechnology building
> Cornell University
> Ithaca, NY 14853
> Email: ok46(a)cornell.edu
>
--
Omry Koren, Ph.D.
Postdoctoral Associate
Department of Microbiology
467 Biotechnology building
Cornell University
Ithaca, NY 14853
Email: ok46(a)cornell.edu
1
0