Reducing the effects of PCR amplification and sequencing artifacts on 16S rRNA-based studies.

Reducing the effects of PCR amplification and sequencing artifacts on 16S rRNA-based studies.
复制标题

DOI:
10.1371/journal.pone.0027310
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Westcott SL
Westcott SL
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Schloss PD;Gevers D;Westcott SL

文献摘要

参考文献

被引文献

相似文献

下一代测序的出现与使用这些方法更好地了解微生物群落的结构和功能在人类,动物和环境健康中的作用的兴趣增长相吻合。然而,使用下一代测序进行16 S rRNA基因序列调查,导致了相当大的争议,测序错误对下游分析的影响。我们分析了分布在90个相同的模拟群落样本中的2.7×106个读数,这些样本是来自21个不同物种的基因组DNA的集合,具有已知的16 S rRNA基因序列;我们观察到平均错误率为0.0060。为了改善这一错误率,我们评估了许多方法来识别坏的序列读数,识别质量差的读数内的区域,并纠正碱基调用,并能够将总体错误率降低到0.0002。PyroNoise算法的实现提供了错误率、序列长度和序列数量的最佳组合。也许比测序错误更成问题的是PCR过程中产生的嵌合体的存在。因为我们知道模拟社区内的真实序列和它们可以形成的嵌合体,我们将8%的原始序列读数鉴定为嵌合体。在对原始序列进行质量过滤并使用Uchime嵌合体检测程序后,总体嵌合率降至1%。不能被检测到的嵌合体在很大程度上负责鉴定虚假的操作分类单位(OTU)和属水平的分类类型。假OTU和OTU类型的数量随着测序工作的增加而增加,这表明应该使用相等数量的序列进行群落比较。最后,我们将我们改进的质量过滤管道应用于几项基准研究,并观察到即使使用我们严格的数据管理管道,也观察到数据生成管道和批次效应中的偏差,这可能会混淆微生物群落数据的解释。
The advent of next generation sequencing has coincided with a growth in interest in using these approaches to better understand the role of the structure and function of the microbial communities in human, animal, and environmental health. Yet, use of next generation sequencing to perform 16S rRNA gene sequence surveys has resulted in considerable controversy surrounding the effects of sequencing errors on downstream analyses. We analyzed 2.7×106 reads distributed among 90 identical mock community samples, which were collections of genomic DNA from 21 different species with known 16S rRNA gene sequences; we observed an average error rate of 0.0060. To improve this error rate, we evaluated numerous methods of identifying bad sequence reads, identifying regions within reads of poor quality, and correcting base calls and were able to reduce the overall error rate to 0.0002. Implementation of the PyroNoise algorithm provided the best combination of error rate, sequence length, and number of sequences. Perhaps more problematic than sequencing errors was the presence of chimeras generated during PCR. Because we knew the true sequences within the mock community and the chimeras they could form, we identified 8% of the raw sequence reads as chimeric. After quality filtering the raw sequences and using the Uchime chimera detection program, the overall chimera rate decreased to 1%. The chimeras that could not be detected were largely responsible for the identification of spurious operational taxonomic units (OTUs) and genus-level phylotypes. The number of spurious OTUs and phylotypes increased with sequencing effort indicating that comparison of communities should be made using an equal number of sequences. Finally, we applied our improved quality-filtering pipeline to several benchmarking studies and observed that even with our stringent data curation pipeline, biases in the data generation pipeline and batch effects were observed that could potentially confound the interpretation of microbial community data.
DOI: 10.1038/nature08821
发表时间: 2010-03-04
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1186/gb-2007-8-7-r143
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Huse SM;Huber JA;Morrison HG;Sogin ML;Welch DM
通讯作者: Welch DM
DOI: 10.1016/s0168-6496(98)00031-2
发表时间: 1998-06-01
影响因子: 4.2
作者:
Hansen, MC;Tolker-Nielsen, T;Molin, S
通讯作者: Molin, S
DOI: 10.1093/bioinformatics/bth226
发表时间: 2004-09-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Huber, T;Faulkner, G;Hugenholtz, P
通讯作者: Hugenholtz, P
DOI: 10.1128/aem.71.12.7724-7736.2005
发表时间: 2005-12-01
影响因子: 4.4
作者:
Ashelford, KE;Chuzhanova, NA;Weightman, AJ
通讯作者: Weightman, AJ