Identification of Bacteriophages Using Deep Representation Model with Pre-training

Identification of Bacteriophages Using Deep Representation Model with Pre-training
复制标题

使用预训练的深度表示模型识别噬菌体

DOI:
10.1101/2021.09.25.461359
复制
发表时间:
2021
期刊:
BioAxiv
影响因子:
--
通讯作者:
Imoto Seiya
Imoto Seiya
中科院分区:
--
文献类型:
--
作者:
Bai Zeheng;Zhang Yao-zhong;Miyano Satoru;Yamaguchi Rui;Uematsu Satoshi;Imoto Seiya

文献摘要

相似文献

噬菌体是一类在细菌和古生菌中进行感染和复制的病毒,在人体中大量存在。要研究微生物与微生物群落的关系,首先要从宏基因组序列中识别出微生物。目前,有两种主要的方法来识别错误:基于数据库的方法和无数据库的方法。基于数据库的方法通常使用大量的序列作为参考;无标记的方法通常使用机器学习和深度学习模型来学习序列的特征。ResultsWe提出了INHERIT,它使用深度表示学习模型来集成基于数据库的方法和无标记的方法,结合了两者的优势。预训练被用作从现有数据库中获取知识表示的替代方式,而BERT风格的深度学习框架保留了无约束方法的优势。我们在第三方基准数据集上将INHERIT与四种现有方法进行了比较。我们的实验表明,INHERIT实现了更好的性能,F1得分为0.9932。此外,我们发现,分别对两个物种进行预训练有助于非对齐深度学习模型做出更准确的预测。可用性和实施INHERIT的代码现在可以在:https://github.com/Celestial-Bai/INHERIT.Supplementary信息补充数据可以在Bioinformaticsonline获得。
MotivationBacteriophages/phages are the viruses that infect and replicate within bacteria and archaea, and rich in human body. To investigate the relationship between phages and microbial communities, the identification of phages from metagenome sequences is the first step. Currently, there are two main methods for identifying phages: database-based (alignment-based) methods and alignment-free methods. Database-based methods typically use a large number of sequences as references; alignment-free methods usually learn the features of the sequences with machine learning and deep learning models.ResultsWe propose INHERIT which uses a deep representation learning model to integrate both database-based and alignment-free methods, combining the strengths of both. Pre-training is used as an alternative way of acquiring knowledge representations from existing databases, while the BERT-style deep learning framework retains the advantage of alignment-free methods. We compare INHERIT with four existing methods on a third-party benchmark dataset. Our experiments show that INHERIT achieves a better performance with the F1-score of 0.9932. In addition, we find that pre-training two species separately helps the non-alignment deep learning model make more accurate predictions.Availability and implementationThe codes of INHERIT are now available in: https://github.com/Celestial-Bai/INHERIT.Supplementary informationSupplementary data are available atBioinformaticsonline.