Removing noise from pyrosequenced amplicons.

Removing noise from pyrosequenced amplicons.
复制标题

DOI:
10.1186/1471-2105-12-38
复制
发表时间:
2011-01-28
期刊:
影响因子:
3
通讯作者:
Turnbaugh PJ
Turnbaugh PJ
中科院分区:
生物学4区
文献类型:
--
作者:
Quince C;Lanzen A;Davenport RJ;Turnbaugh PJ

文献摘要

被引文献

相似文献

在许多环境基因组学应用中,来自不同样本的同源DNA区域首先通过PCR扩增,然后测序。下一代测序技术,454焦磷酸测序,使得PCR扩增子的读取数比以往任何时候都大得多。这已经彻底改变了微生物多样性的研究,因为现在有可能对一个群落中的大部分16S rRNA基因进行测序。然而,越来越多的人认识到,由于大量的读取数和缺乏共识序列,在这些数据中区分噪声和真正的序列多样性至关重要。否则,这将导致对现有类型或操作分类单位(otu)数量的夸大估计。三种错误来源是重要的:测序错误,PCR单碱基替换和PCR嵌合体。我们提出了AmpliconNoise, PyroNoise算法的发展,能够分别去除454个测序错误和PCR单碱基错误。我们还介绍了一种新的嵌合体去除程序,Perseus,利用与焦磷酸测序数据相关的序列丰度。我们使用数据集,其中已知多样性的样本已经被放大和测序,以量化每种误差来源对OTU膨胀的影响,并验证这些算法。AmpliconNoise优于其他算法,大大降低了GS FLX和最新Titanium协议的每基错误率。所有这三个误差来源都会导致多样性估计的膨胀。特别是,嵌合体的形成具有迄今为止尚未认识到的重要性,其重要性因扩增方案而异。我们表明,AmpliconNoise可以准确估计OTU数。同样重要的是,AmpliconNoise即使在低序列差异下也能产生正确的otu。我们证明了珀尔修斯有很高的灵敏度,能够发现99%的嵌合体,当这些嵌合体出现在高频率时,这是至关重要的。AmpliconNoise紧随其后的Perseus是一个非常有效的去除噪声的管道。此外,算法背后的原理,使用期望最大化(EM)对真序列的推断,以及将嵌合体检测作为分类或“监督学习”问题的处理,将同样适用于新测序技术。
In many environmental genomics applications a homologous region of DNA from a diverse sample is first amplified by PCR and then sequenced. The next generation sequencing technology, 454 pyrosequencing, has allowed much larger read numbers from PCR amplicons than ever before. This has revolutionised the study of microbial diversity as it is now possible to sequence a substantial fraction of the 16S rRNA genes in a community. However, there is a growing realisation that because of the large read numbers and the lack of consensus sequences it is vital to distinguish noise from true sequence diversity in this data. Otherwise this leads to inflated estimates of the number of types or operational taxonomic units (OTUs) present. Three sources of error are important: sequencing error, PCR single base substitutions and PCR chimeras. We present AmpliconNoise, a development of the PyroNoise algorithm that is capable of separately removing 454 sequencing errors and PCR single base errors. We also introduce a novel chimera removal program, Perseus, that exploits the sequence abundances associated with pyrosequencing data. We use data sets where samples of known diversity have been amplified and sequenced to quantify the effect of each of the sources of error on OTU inflation and to validate these algorithms. AmpliconNoise outperforms alternative algorithms substantially reducing per base error rates for both the GS FLX and latest Titanium protocol. All three sources of error lead to inflation of diversity estimates. In particular, chimera formation has a hitherto unrealised importance which varies according to amplification protocol. We show that AmpliconNoise allows accurate estimates of OTU number. Just as importantly AmpliconNoise generates the right OTUs even at low sequence differences. We demonstrate that Perseus has very high sensitivity, able to find 99% of chimeras, which is critical when these are present at high frequencies. AmpliconNoise followed by Perseus is a very effective pipeline for the removal of noise. In addition the principles behind the algorithms, the inference of true sequences using Expectation-Maximization (EM), and the treatment of chimera detection as a classification or 'supervised learning' problem, will be equally applicable to new sequencing technologies as they appear.