Hi, Sorting by '# reads hit' seems more suitable as it will give you the genes in the reference organism which are most abundant in the metagenomic dataset. It is based purely on the similarity information and the numbers in each row will include effects due to number of organisms with similarity and gene copy number. While this will give you the more abundant genes, I don't know how well this will correlate with the importance of the function for the survival in an environment. The numbers next to the organism name in the reference organism listing represents the number of hits against that organism, sorted in descending order so that the organism at the top generally displays greater similarity to the metagenomic data than the one at the bottom, especially when the difference in the number of hits is large. You are welcome, glad to be of help, Mark ----- Original Message -----
From: "Joanne Wyglinski, Ms" <[email protected]> To: "Mark DSouza" <[email protected]> Sent: Thursday, August 25, 2011 11:57:24 AM Subject: RE: [mg-rast] # reads hit Dear Mark,
Thank you for your response, it is greatly appreciated.
I am looking for which genes in relevant organisms in the reference set (ie. thermophiles) are the most abundant in my sample so that I can figure out which genes play the most important role in terms of how these organisms survive or thrive in the environment that they are in. Do the # reads hit take into account the number of organisms with similarity or gene copy number?
Please let me know if what I am doing makes sense or if I should be organizing according to evalue or % identity.
My other question was what the numbers next to the organism name in the reference set pull down menu stand for. Many organisms have the same number, so I was wondering why that is.
Thanks again for your help.
Best regards, Joanne ________________________________________ From: Mark DSouza [[email protected]] Sent: August 25, 2011 12:37 PM To: Joanne Wyglinski, Ms Cc: [email protected] Subject: Re: [mg-rast] # reads hit
Hi,
The '# reads hit' refers to the number of reads in the metagenomic sample that were found to have sequence similarity to a gene in the reference genome, sorting this column in descending order puts the genes which hit the largest number of reads at the top of the table. This would give you a sense of the content of the metagenomic dataset (as compared to the reference genome) while organizing by average evalue or %identity will highlight other genes based on similarity to genes in the reference genome. If you let us know what you are looking for in your metagenomic dataset we will try and provide a more detailed answer.
-- Regards, Mark D'Souza -- for the MG-RAST team
[email protected] http://metagenomics.anl.gov/
If you need to respond to this email please address it to the mailing-list [email protected]
----- Original Message -----
From: "Joanne Wyglinski, Ms" <[email protected]> To: [email protected] Sent: Thursday, August 25, 2011 8:55:50 AM Subject: Re: [mg-rast] # reads hit Hi,
I have a question regarding metagenomic analysis on your website.
For the organism recruitment plot, the MG-RAST website states that you can organize according to "# reads hit".
What does this mean?
Also, if I am looking for the most significant data, should I be organizing according to "# reads hit" or something like avg evalue or % identity.
Thanks for your time in advance. Your help is greatly appreciated.
Best regards, Joanne Wyglinski
participants (1)
-
Mark DSouza