The Promises and Pitfalls of Machine Learning for Detecting Viruses in Aquatic Metagenomes

The Promises and Pitfalls of Machine Learning for Detecting Viruses in Aquatic Metagenomes
复制标题

DOI:
10.3389/fmicb.2019.00806
复制
发表时间:
2019-04-16
影响因子:
5.2
通讯作者:
Hurwitz, Bonnie L.
Hurwitz, Bonnie L.
中科院分区:
生物学2区
文献类型:
--
作者:
Ponsero, Alise J.;Hurwitz, Bonnie L.

文献摘要

被引文献

相似文献

允许在宿主相关和环境宏基因组中鉴定病毒序列的工具可以更好地理解病毒及其宿主的遗传学和生态。最近,使用机器学习方法使用K-MER序列特征来区分病毒和细菌信号的新方法,以鉴定宏基因组中的病毒重叠群。这些基于内容的方法的承诺是发现新病毒的能力,没有或几个已知的亲戚。在此观点论文中,我们研究了基于内容的机器学习工具Virfinder在水生元基因组中识别病毒序列的使用,并探讨了使用针对海洋元基因组的以生态系统为中心的模型的可能性。我们讨论了训练集组成对刀具性能的影响,以及在元基因组中检索低丰度病毒序列的当前局限性。我们确定可能是由现实世界中的机器学习方法引起的潜在偏见,并提出了克服它们的可能途径。
Tools allowing for the identification of viral sequences in host-associated and environmental metagenomes allows for a better understanding of the genetics and ecology of viruses and their hosts. Recently, new approaches using machine learning methods to distinguish viral from bacterial signal using k-mer sequence signatures were published for identifying viral contigs in metagenomes. The promise of these content-based approaches is the ability to discover new viruses, with no or few known relatives. In this perspective paper, we examine the use of the content-based machine learning tool VirFinder for the identification of viral sequences in aquatic metagenomes and explore the possibility of using ecosystem-focused models targeted to marine metagenomes. We discuss the impact of the training set composition on the tool performance and the current limitation for the retrieval of low abundance viral sequences in metagenomes. We identify potential biases that could arise from machine learning approaches for viral hunting in real-world datasets and suggest possible avenues to overcome them.