RAIphy: phylogenetic classification of metagenomics samples using iterative refinement of relative abundance index profiles.

RAIphy: phylogenetic classification of metagenomics samples using iterative refinement of relative abundance index profiles.
复制标题

DOI:
10.1186/1471-2105-12-41
复制
发表时间:
2011-01-31
期刊:
影响因子:
3
通讯作者:
Sayood K
Sayood K
中科院分区:
生物学4区
文献类型:
--
作者:
Nalbantoglu OU;Way SF;Hinrichs SH;Sayood K

文献摘要

参考文献

被引文献

相似文献

元基因组的计算分析需要对从环境样本的DNA读数组装的基因组重叠群进行分类分配。由于微生物群的多样性,所获得的组合体的长度可以从几百个碱基对到几百千个碱基不等。目前的分类算法为来自与注释基因组有近亲的生物体的长重叠群或短片段提供了准确的分类。由于微生物群的复杂性和现有注释基因组的匮乏,这些都是元基因组分析的重大限制。我们提出了一种稳健的分类方法RAIphy,它使用了一种新的序列相似性度量,并对分类模型和函数进行了迭代求精,有效地避免了这些限制。我们已经用合成的元基因组数据测试了RAIphy,数据范围从100个基点到50个基点。在100bp-1000bp的序列读取范围内,RAIphy的灵敏度范围在38%-81%之间,优于目前流行的基于成分的方法在该范围内的读取。与计算量更大的序列相似性方法的比较表明,RAIphy的性能更具竞争力,同时速度更快。将相对较长的重叠群的敏感性-特异性特征与PHYLOPITIA和TACOA算法进行了比较。在不同的分支级别上,RAIphy比这些算法执行得更好。对于酸性矿山废水(AMD)元基因组,RAIphy能够在分类上比目前可用的方法Phymm和Megan更准确地结合序列读取集,并且在三个测试中有两个比计算更密集的方法Phymm BL更准确。通过引入相对丰度指数度量和迭代分类方法,我们提出了一种分类分类算法,该算法对从元基因组数据组装的大范围DNA重叠群长度具有竞争性。由于其快速、简单和准确,RAIphy可以成功地用于从环境样本获得的广泛的元基因组数据的入库过程。
Computational analysis of metagenomes requires the taxonomical assignment of the genome contigs assembled from DNA reads of environmental samples. Because of the diverse nature of microbiomes, the length of the assemblies obtained can vary between a few hundred bp to a few hundred Kbp. Current taxonomic classification algorithms provide accurate classification for long contigs or for short fragments from organisms that have close relatives with annotated genomes. These are significant limitations for metagenome analysis because of the complexity of microbiomes and the paucity of existing annotated genomes. We propose a robust taxonomic classification method, RAIphy, that uses a novel sequence similarity metric with iterative refinement of taxonomic models and functions effectively without these limitations. We have tested RAIphy with synthetic metagenomics data ranging between 100 bp to 50 Kbp. Within a sequence read range of 100 bp-1000 bp, the sensitivity of RAIphy ranges between 38%-81% outperforming the currently popular composition-based methods for reads in this range. Comparison with computationally more intensive sequence similarity methods shows that RAIphy performs competitively while being significantly faster. The sensitivity-specificity characteristics for relatively longer contigs were compared with the PhyloPythia and TACOA algorithms. RAIphy performs better than these algorithms at varying clade-levels. For an acid mine drainage (AMD) metagenome, RAIphy was able to taxonomically bin the sequence read set more accurately than the currently available methods, Phymm and MEGAN, and more accurately in two out of three tests than the much more computationally intensive method, PhymmBL. With the introduction of the relative abundance index metric and an iterative classification method, we propose a taxonomic classification algorithm that performs competitively for a large range of DNA contig lengths assembled from metagenome data. Because of its speed, simplicity, and accuracy RAIphy can be successfully used in the binning process for a broad range of metagenomic data obtained from environmental samples.
DOI: 10.1080/07391102.1986.10507643
发表时间: 1986-08-01
影响因子: 4.4
作者:
BRENDEL, V;BECKMANN, JS;TRIFONOV, EN
通讯作者: TRIFONOV, EN
DOI: 10.1186/1471-2105-10-56
发表时间: 2009-02-11
期刊: BMC bioinformatics
影响因子: 3
作者:
Diaz NN;Krause L;Goesmann A;Niehaus K;Nattkemper TW
通讯作者: Nattkemper TW
DOI: 10.1038/nmeth.1358
发表时间: 2009-09
期刊: NATURE METHODS
影响因子: 48
作者:
Brady, Arthur;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1093/nar/gkl842
发表时间: 2007-01
影响因子: 14.9
作者:
Pruitt KD;Tatusova T;Maglott DR
通讯作者: Maglott DR
DOI: 10.1101/gr.5969107
发表时间: 2007-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Huson, Daniel H.;Auch, Alexander F.;Schuster, Stephan C.
通讯作者: Schuster, Stephan C.