Classifying short genomic fragments from novel lineages using composition and homology.

Classifying short genomic fragments from novel lineages using composition and homology.
复制标题

DOI:
10.1186/1471-2105-12-328
复制
发表时间:
2011-08-09
期刊:
影响因子:
3
通讯作者:
Beiko RG
Beiko RG
中科院分区:
生物学4区
文献类型:
--
作者:
Parks DH;MacDonald NJ;Beiko RG

文献摘要

参考文献

被引文献

相似文献

直接从环境中回收的DNA片段的分类属性的分配是宏基因组数据分析的重要步骤。可以使用等级特异性分类器或等级灵活分类器进行分类,等级特异性分类器将读数分配给来自预定水平的分类标签,例如命名的物种或菌株,等级灵活分类器为数据集中的每个序列选择适当的分类等级。排名的选择通常取决于给定序列的最佳模型和在一组接近最佳模型中看到的分类组的宽度。基于同源性(例如,LCA)和基于成分的(例如,PhyloPythia、TACOA)已经提出了等级灵活的分类器,但目前还没有同时利用同源性和组合性的混合方法。我们首先开发了一个混合的,基于BLAST和朴素贝叶斯(NB)的特定于秩的分类器,它具有可比的准确性和比目前最好的方法PhymmBL更快的运行时间。通过替换LCA的BLAST或允许包含次优NB模型,我们得到一个排名灵活的分类器。这种混合分类器优于建立的排名灵活的方法上模拟的宏基因组片段的长度为200 bp至1000 bp,并能够分配分类属性的序列的子集,很少有错误分类。然后,我们证明了不同的分类器的性能增强的生物除磷宏基因组,说明了排名灵活的分类器的优势时,代表性的基因组是不存在的参考基因组。冰川冰宏基因组的应用程序表明,类似的分类配置文件中获得的一组分类越来越保守的分类。我们基于NB的分类方案比目前最好的基于组合的算法Phymm更快,同时提供同样准确的预测。NB的秩可变变体,我们称之为ε-NB,是LCA的补充,可以与它结合,产生非常高置信度的保守预测集。LCA和ε-NB的简单参数化允许调整更多预测和更高精度之间的平衡,允许用户考虑下游分析对错误分类或未分类序列的敏感性。
The assignment of taxonomic attributions to DNA fragments recovered directly from the environment is a vital step in metagenomic data analysis. Assignments can be made using rank-specific classifiers, which assign reads to taxonomic labels from a predetermined level such as named species or strain, or rank-flexible classifiers, which choose an appropriate taxonomic rank for each sequence in a data set. The choice of rank typically depends on the optimal model for a given sequence and on the breadth of taxonomic groups seen in a set of close-to-optimal models. Homology-based (e.g., LCA) and composition-based (e.g., PhyloPythia, TACOA) rank-flexible classifiers have been proposed, but there is at present no hybrid approach that utilizes both homology and composition. We first develop a hybrid, rank-specific classifier based on BLAST and Naïve Bayes (NB) that has comparable accuracy and a faster running time than the current best approach, PhymmBL. By substituting LCA for BLAST or allowing the inclusion of suboptimal NB models, we obtain a rank-flexible classifier. This hybrid classifier outperforms established rank-flexible approaches on simulated metagenomic fragments of length 200 bp to 1000 bp and is able to assign taxonomic attributions to a subset of sequences with few misclassifications. We then demonstrate the performance of different classifiers on an enhanced biological phosphorous removal metagenome, illustrating the advantages of rank-flexible classifiers when representative genomes are absent from the set of reference genomes. Application to a glacier ice metagenome demonstrates that similar taxonomic profiles are obtained across a set of classifiers which are increasingly conservative in their classification. Our NB-based classification scheme is faster than the current best composition-based algorithm, Phymm, while providing equally accurate predictions. The rank-flexible variant of NB, which we term ε-NB, is complementary to LCA and can be combined with it to yield conservative prediction sets of very high confidence. The simple parameterization of LCA and ε-NB allows for tuning of the balance between more predictions and increased precision, allowing the user to account for the sensitivity of downstream analyses to misclassified or unclassified sequences.
Metasim:用于基因组学和元基因组学的测序模拟器。
DOI: 10.1371/journal.pone.0003373
发表时间: 2008-10-08
期刊: PLOS ONE
影响因子: 3.7
作者:
Richter, Daniel C.;Ott, Felix;Auch, Alexander F.;Schmid, Ramona;Huson, Daniel H.
通讯作者: Huson, Daniel H.
DOI: 10.1038/nature08821
发表时间: 2010-03-04
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1093/bioinformatics/btq041
发表时间: 2010-03-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Parks, Donovan H.;Beiko, Robert G.
通讯作者: Beiko, Robert G.
DOI: 10.2307/1403452
发表时间: 2001-12-01
影响因子: 2
作者:
Hand, DJ;Yu, KM
通讯作者: Yu, KM
DOI: 10.1093/nar/gkn721
发表时间: 2009-01
影响因子: 14.9
作者:
Pruitt KD;Tatusova T;Klimke W;Maglott DR
通讯作者: Maglott DR