Re: [mg-rast] Fwd: Re: Analysis question
Hi Analysis team, please see below. I wanted to ask what the status of this is. I checked my data set and it seems that you solved the problem since the numbers for Speleothem B went down. But I wanted to confirm with you. Thanks, Antje Quoting Mark D'Souza <[email protected]>:
Hi,
Your explanation was clear, I had mis-read your earlier email.
The numbers for SpeleothemB do seem unusual, we are rerunning one of the pipeline stages which may be responsible, I will get back to you when this is done.
Regards, Mark
On Jul 11, 2011, at 11:34 AM, [email protected] wrote:
Hi Analysis team, thanks for your fast response. The MG-Rast ID's for the three samples are: 4466357.3 (Speleothem A); 4466359.3 (Wet Rock); 4466556.3 (Speleothem B) 4466556.3 (Speleothem B) is the sample I NAMED IN MY FIRST EMAIL SAMPLE A.
The number of counts for sample A were significantly higher. For example when I compared the three samples in the organisms classification mode using the SEED database (max. e-value cutoff: 1e-5, min. identity cutoff: 50%, min. alignment cutoff: 50, raw data) than I got for the bacteria bar (barchart graph):
MG-RAST ID 4466556.3 - 2,454,565 counts MG-RAST ID 4466359.3 - 488,597 counts MG RAST ID 4466556.3 - 447,833 counts. For the other categories (unassigned, Archaea...) the counts of MG-RAST ID 4466556.3 were also much higher compared to the other two samples. When I chose another database to compare it looked similar. For example when I chose Swiss Prot (same settings like before fore SEED), than I got for bacteria: MG-RAST ID 4466556.3 - 567,335 counts MG-RAST ID 4466359.3 - 81,623 counts MG-RAST ID 4466357.3 - 71,885 counts.
Please let me know, if you have any questions regarding my explanation.
Regards,
Antje
-- Antje Legatzki Dept Soil Water and Environmental Science Shantz Bldg. Rm 429 University of Arizona Tucson, AZ 85721 520-621-9759
----- Forwarded message from [email protected] ----- Date: Mon, 11 Jul 2011 10:33:23 -0500 From: Mark D'Souza Reply-To: Mark D'Souza Subject: Re: [mg-rast] Analysis question To: [email protected]
Hi,
Yes, once you redraw the barchart using raw values, you will see the number of hits in your metagenomic sample.
Each read can have multiple hits against the database, so it is possible that the number of hits exceeds the number of sequences, can you send us the MG-RAST IDs for the samples A,B &C, we can check that this is the case and look into the discrepancy in the counts for hits against bacteria.
To clarify, you say "we get with this analysis about 2,450,000 counts for bacteria (it is similar with the other databases)" and "Sample A has always much more assigned sequences than sample B and C". Is the number for A significantly higher than B and C?
Regards, Mark D'Souza -- for the MG-RAST team
[email protected] http://metagenomics.anl.gov/
If you need to respond to this email please address it to the mailing-list [email protected]
On Jul 8, 2011, at 7:42 PM, [email protected] wrote:
Hi MG-RAST support team, we uploaded several metagenomic data sets to MG-RAST, I will call them here A,B, and C. I have a couple of questions regarding the organisms classification of one of our data sets with the barchart option (raw data). First: Do I understand it right that if I move the cursor over the regenerated barchart, that the number in the brackets are the counts for the hits?
For the second question I have to explain further: As an example I generated a barchart graph using the SEED database, max. e-value 1e-5, min % identity cut off 50%, min alignment length 50. We have about 780,000 sequences after post QC and we get with this analysis about 2,450,000 counts for bacteria (it is similar with the other databases). These are much more hits than we have sequences. Is this real? Do you have any idea how this can be?
And third: We have two other metagenomic datasets (B, C) from the same environment (from other sample sites). Those gave us post QC about 750,000 and 780,000 sequences, respectively. If we use the same settings like for sample A for the taxonomic classification, than we got 450,000 or 490,000 assigned sequences for bacteria (bacteria are in all 3 samples the dominant group). Sample A has always much more assigned sequences than sample B and C. The sample sites are not so different that I think the difference is real. Do you have any explanation for this? If this seems not normal, do you have any suggestion what we should do (for example run the samples again through the pipeline?
Thanks,
Antje
-- Antje Legatzki Dept Soil Water and Environmental Science Shantz Bldg. Rm 429 University of Arizona Tucson, AZ 85721 520-621-9759
----- End forwarded message -----
participants (1)
-
legatzkiļ¼ email.arizona.edu