ProViDE: A software tool for accurate estimation of viral diversity in metagenomic samples

ProViDE: A software tool for accurate estimation of viral diversity in metagenomic samples
复制标题

DOI:
10.6026/97320630006091
复制
发表时间:
2011-01-01
期刊:
影响因子:
1.9
通讯作者:
Mande, Sharmila Shekhar
Mande, Sharmila Shekhar
中科院分区:
其他
文献类型:
--
作者:
Ghosh, Tarini Shankar;Mohammed, Monzoorul Haque;Mande, Sharmila Shekhar

文献摘要

被引文献

相似文献

鉴于病毒领域缺乏通用标记基因,研究人员通常使用BLAST(具有严格的e值)对病毒宏基因组序列进行分类分类。由于大多数宏基因组序列来自迄今未知的病毒群,使用严格的e值导致大多数序列仍未分类。此外,使用不太严格的e值会导致大量不正确的分类分配。SOrt-ITEMS算法提供了一种解决上述问题的方法。基于比对参数,SOrt-ITEMS遵循一个复杂的工作流程来分配来自迄今未知的古细菌/细菌基因组的reads。在SOrt-ITEMS中,通过观察属于细菌和古细菌王国的不同分类类群内部和之间的序列差异模式来产生比对参数阈值。然而,病毒王国内的许多分类类群缺乏典型的林奈式分类等级。在本文中,我们提出了提供(病毒多样性估计程序),这是一种算法,它使用一组定制的比对参数阈值,特别适合病毒宏基因组序列。这些阈值捕获了序列分化的模式和在病毒王国的不同分类群内/之间观察到的不统一的分类等级。验证结果表明,提供的“正确”分配百分比比广泛使用的基于相似性的方法MEGAN高1.7到3倍。提供的错误分类率约为3%至19% (MEGAN的错误分类率为5%至42%),表明分配准确性明显更好。
Given the absence of universal marker genes in the viral kingdom, researchers typically use BLAST (with stringent E-values) for taxonomic classification of viral metagenomic sequences. Since majority of metagenomic sequences originate from hitherto unknown viral groups, using stringent e-values results in most sequences remaining unclassified. Furthermore, using less stringent e-values results in a high number of incorrect taxonomic assignments. The SOrt-ITEMS algorithm provides an approach to address the above issues. Based on alignment parameters, SOrt-ITEMS follows an elaborate work-flow for assigning reads originating from hitherto unknown archaeal/bacterial genomes. In SOrt-ITEMS, alignment parameter thresholds were generated by observing patterns of sequence divergence within and across various taxonomic groups belonging to bacterial and archaeal kingdoms. However, many taxonomic groups within the viral kingdom lack a typical Linnean-like taxonomic hierarchy. In this paper, we present ProViDE (Program for Viral Diversity Estimation), an algorithm that uses a customized set of alignment parameter thresholds, specifically suited for viral metagenomic sequences. These thresholds capture the pattern of sequence divergence and the non-uniform taxonomic hierarchy observed within/across various taxonomic groups of the viral kingdom. Validation results indicate that the percentage of 'correct' assignments by ProViDE is around 1.7 to 3 times higher than that by the widely used similarity based method MEGAN. The misclassification rate of ProViDE is around 3 to 19% (as compared to 5 to 42% by MEGAN) indicating significantly better assignment accuracy.