An improved deep learning method for predicting DNA-binding proteins based on contextual features in amino acid sequences

An improved deep learning method for predicting DNA-binding proteins based on contextual features in amino acid sequences
复制标题

DOI:
10.1371/journal.pone.0225317
复制
发表时间:
2019-11-14
期刊:
影响因子:
3.7
通讯作者:
Wang, Haiou
Wang, Haiou
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hu, Siquan;Ma, Ruixiong;Wang, Haiou

文献摘要

被引文献

相似文献

随着已知蛋白质数量的不断增加,如何准确鉴定DNA结合蛋白已成为一个重大的生物学挑战。目前,已经提出了各种计算方法来识别仅从氨基酸序列的DNA结合蛋白,如SVM,DNABP和CNN-RNN。然而,这些方法没有考虑氨基酸序列中的背景,这使得它们难以充分捕获序列特征。在这项研究中,提出了一种协调双向长期记忆递归神经网络和卷积神经网络(称为CNN-BiLSTM)的新方法来识别DNA结合蛋白。CNN-BiLSTM模型可以探索氨基酸序列的潜在上下文关系,并获得比传统模型更多的特征。实验结果表明,CNN-BiLSTM的验证集预测精度比SVM高96.5%-7.8%,比DNABP高9.6%,比CNN-RNN高3.7%。在UniProt提供的20,000个不参与模型训练的独立样本上测试后,CNN-BiLSTM的准确率比SVM高94.5%-12%,比DNABP高4.9%,比CNN-RNN高4%。我们可视化并比较了CNN-BiLSTM和CNN-RNN的模型训练过程,发现前者能够更好地从训练数据集泛化,表明CNN-BiLSTM对蛋白质序列具有更广泛的适应性。在测试集上,CNN-BiLSTM具有更好的可信度,因为它的预测分数比CNN-RNN更接近样本标签。因此,提出的CNN-BiLSTM是一种更强大的识别DNA结合蛋白的方法。
As the number of known proteins has expanded, how to accurately identify DNA binding proteins has become a significant biological challenge. At present, various computational methods have been proposed to recognize DNA-binding proteins from only amino acid sequences, such as SVM, DNABP and CNN-RNN. However, these methods do not consider the context in amino acid sequences, which makes it difficult for them to adequately capture sequence features. In this study, a new method that coordinates a bidirectional long-term memory recurrent neural network and a convolutional neural network, called CNN-BiLSTM, is proposed to identify DNA binding proteins. The CNN-BiLSTM model can explore the potential contextual relationships of amino acid sequences and obtain more features than can traditional models. The experimental results show that the CNN-BiLSTM achieves a validation set prediction accuracy of 96.5%-7.8% higher than that of SVM, 9.6% higher than that of DNABP and 3.7% higher than that of CNN-RNN. After testing on 20,000 independent samples provided by UniProt that were not involved in model training, the accuracy of CNN-BiLSTM reached 94.5%-12% higher than that of SVM, 4.9% higher than that of DNABP and 4% higher than that of CNN-RNN. We visualized and compared the model training process of CNN-BiLSTM with that of CNN-RNN and found that the former is capable of better generalization from the training dataset, showing that CNN-BiLSTM has a wider range of adaptations to protein sequences. On the test set, CNN-BiLSTM has better credibility because its predicted scores are closer to the sample labels than are those of CNN-RNN. Therefore, the proposed CNN-BiLSTM is a more powerful method for identifying DNA-binding proteins.