Identifying SNAREs by Incorporating Deep Learning Architecture and Amino Acid Embedding Representation

Identifying SNAREs by Incorporating Deep Learning Architecture and Amino Acid Embedding Representation
复制标题

DOI:
10.3389/fphys.2019.01501
复制
发表时间:
2019-12-10
影响因子:
4
通讯作者:
Huynh, Tuan-Tu
Huynh, Tuan-Tu
中科院分区:
医学2区
文献类型:
--
作者:
Nguyen Quoc Khanh Lee;Huynh, Tuan-Tu

文献摘要

被引文献

相似文献

SNARE(可溶性N-乙基马来酰亚胺敏感因子激活蛋白受体)是一组对细胞膜融合和神经递质胞吐至关重要的蛋白质。它们在广泛的细胞过程中发挥重要作用,包括细胞生长、胞质分裂和突触传递,以促进真核生物中的细胞膜整合。许多研究表明,SNARE蛋白与许多人类疾病,特别是癌症有关。因此,确定它们的功能是科学家更好地了解癌症疾病以及设计治疗药物靶点的一个具有挑战性的问题。我们使用fastText描述了基于氨基酸嵌入的每个蛋白质序列,fastText是一种在其领域表现良好的自然语言处理模型。由于每个蛋白质序列类似于一个包含不同单词的句子,因此将语言模型应用于蛋白质序列具有挑战性和前景。生成后,氨基酸嵌入特征被输入深度学习算法进行预测。我们的模型结合了fastText模型和深度卷积神经网络,可以识别SNARE蛋白,独立测试准确率为92.8%,灵敏度为88.5%,特异性为97%,马修斯相关系数(MCC)为0.86。我们的性能结果上级最先进的预测器(SNARE-CNN)。本研究为生物学家识别SNARE提供了一种可靠的方法,为将fastText词嵌入模型应用于生物信息学,特别是蛋白质序列预测奠定了基础。
SNAREs (soluble N-ethylmaleimide-sensitive factor activating protein receptors) are a group of proteins that are crucial for membrane fusion and exocytosis of neurotransmitters from the cell. They play an important role in a broad range of cell processes, including cell growth, cytokinesis, and synaptic transmission, to promote cell membrane integration in eukaryotes. Many studies determined that SNARE proteins have been associated with a lot of human diseases, especially in cancer. Therefore, identifying their functions is a challenging problem for scientists to better understand the cancer disease as well as design the drug targets for treatment. We described each protein sequence based on the amino acid embeddings using fastText, which is a natural language processing model performing well in its field. Because each protein sequence is similar to a sentence with different words, applying language model into protein sequence is challenging and promising. After generating, the amino acid embedding features were fed into a deep learning algorithm for prediction. Our model which combines fastText model and deep convolutional neural networks could identify SNARE proteins with an independent test accuracy of 92.8%, sensitivity of 88.5%, specificity of 97%, and Matthews correlation coefficient (MCC) of 0.86. Our performance results were superior to the state-of-the-art predictor (SNARE-CNN). We suggest this study as a reliable method for biologists for SNARE identification and it serves a basis for applying fastText word embedding model into bioinformatics, especially in protein sequencing prediction.