MARVEL, a Tool for Prediction of Bacteriophage Sequences in Metagenomic Bins

MARVEL, a Tool for Prediction of Bacteriophage Sequences in Metagenomic Bins
复制标题

DOI:
10.3389/fgene.2018.00304
复制
发表时间:
2018-08-07
影响因子:
3.7
通讯作者:
Setubal, Joao C.
Setubal, Joao C.
中科院分区:
生物学3区
文献类型:
--
作者:
Amgarten, Deyvid;Braga, Lucas P. P.;Setubal, Joao C.

文献摘要

被引文献

相似文献

在这里,我们提出了MARVEL,预测宏基因组箱中双链DNA噬菌体序列的工具。MARVEL使用随机森林机器学习方法。我们在具有1,247个噬菌体和1,029个细菌基因组的数据集上训练了该程序,并在具有335个细菌和177个噬菌体基因组的数据集上对其进行了测试。我们表明,从重叠群序列中提取的三个简单的基因组特征足以实现从噬菌体序列中分离细菌的良好性能:基因密度,链转移和病毒蛋白质数据库的显著命中分数。我们将MARVEL的性能与VirSorter和Virginia(两种流行的病毒序列预测程序)的性能进行了比较。我们的研究结果表明,所有三个程序具有可比的特异性,但MARVEL实现更好的性能召回(敏感性)措施。这意味着MARVEL应该能够在宏基因组箱中识别比迄今为止可能的更多的噬菌体序列。在一个简单的测试与真实的数据,主要包含细菌序列,MARVEL分类58的209箱作为噬菌体基因组;其他证据表明,这58箱中的57个是新的噬菌体序列。
Here we present MARVEL, a tool for prediction of double-stranded DNA bacteriophage sequences in metagenomic bins. MARVEL uses a random forest machine learning approach. We trained the program on a dataset with 1,247 phage and 1,029 bacterial genomes, and tested it on a dataset with 335 bacterial and 177 phage genomes. We show that three simple genomic features extracted from contig sequences were sufficient to achieve a good performance in separating bacterial from phage sequences: gene density, strand shifts, and fraction of significant hits to a viral protein database. We compared the performance of MARVEL to that of VirSorter and VirFinder, two popular programs for predicting viral sequences. Our results show that all three programs have comparable specificity, but MARVEL achieves much better performance on the recall (sensitivity) measure. This means that MARVEL should be able to identify many more phage sequences in metagenomic bins than heretofore has been possible. In a simple test with real data, containing mostly bacterial sequences, MARVEL classified 58 out of 209 bins as phage genomes; other evidence suggests that 57 of these 58 bins are novel phage sequences.