iACVP: markedly enhanced identification of anti-coronavirus peptides using a dataset-specific word2vec model

iACVP: markedly enhanced identification of anti-coronavirus peptides using a dataset-specific word2vec model
复制标题

DOI:
10.1093/bib/bbac265
复制
发表时间:
2022-07-01
影响因子:
9.5
通讯作者:
Manavalan, Balachandran
Manavalan, Balachandran
中科院分区:
生物学2区
文献类型:
--
作者:
Kurata, Hiroyuki;Tsukiyama, Sho;Manavalan, Balachandran

文献摘要

被引文献

相似文献

COVID-19大流行在全球造成数百万人死亡。因此,开发抗冠状病毒药物迫在眉睫。与常规的非多肽药物不同,抗病毒多肽药物具有高度特异性,易于合成和修饰,不易产生耐药性。为了减少筛选数千种肽并测定其抗病毒活性所涉及的时间和费用,需要用于识别抗冠状病毒肽(ACVPs)的计算预测因子。然而,尽管已经发现了相对大量的抗病毒肽(AVPs),但实验验证的ACVP样品很少。在本研究中,我们尝试使用AVP数据集和一小部分acvp集合来预测acvp。利用常规特征、二进制轮廓和嵌入词的word2vec (W2V),我们系统地探索了五种不同的机器学习方法:变压器、卷积神经网络、双向长短期记忆、随机森林(RF)和支持向量机。通过穷举搜索,我们发现具有W2V的RF分类器在不同的数据集上一致地取得了更好的性能。两个主要控制因素是:(i)从训练数据集和独立测试数据集生成特定数据集的W2V字典,而不是广泛使用的通用UniProt蛋白质组;(ii)进行系统搜索并确定W2V的最佳k-mer值,这提供了更好的阳性和阴性样本区分。因此,与现有的最先进的方法相比,我们提出的方法(称为iACVP)始终提供更好的预测性能。为了帮助实验人员确定假定的acvp,我们将我们的模型实现为一个web服务器,可通过以下链接访问:http://kurata35.bio.kyutech.ac.jp/iACVP。
The COVID-19 pandemic caused several million deaths worldwide. Development of anti-coronavirus drugs is thus urgent. Unlike conventional non-peptide drugs, antiviral peptide drugs are highly specific, easy to synthesize and modify, and not highly susceptible to drug resistance. To reduce the time and expense involved in screening thousands of peptides and assaying their antiviral activity, computational predictors for identifying anti-coronavirus peptides (ACVPs) are needed. However, few experimentally verified ACVP samples are available, even though a relatively large number of antiviral peptides (AVPs) have been discovered. In this study, we attempted to predict ACVPs using an AVP dataset and a small collection of ACVPs. Using conventional features, a binary profile and a word-embedding word2vec (W2V), we systematically explored five different machine learning methods: Transformer, Convolutional Neural Network, bidirectional Long Short-Term Memory, Random Forest (RF) and Support Vector Machine. Via exhaustive searches, we found that the RF classifier with W2V consistently achieved better performance on different datasets. The two main controlling factors were: (i) the dataset-specific W2V dictionary was generated from the training and independent test datasets instead of the widely used general UniProt proteome and (ii) a systematic search was conducted and determined the optimal k-mer value in W2V, which provides greater discrimination between positive and negative samples. Therefore, our proposed method, named iACVP, consistently provides better prediction performance compared with existing state-of-the-art methods. To assist experimentalists in identifying putative ACVPs, we implemented our model as a web server accessible via the following link: http://kurata35.bio.kyutech.ac.jp/iACVP.