Using deep neural networks and biological subwords to detect protein S-sulfenylation sites

Using deep neural networks and biological subwords to detect protein S-sulfenylation sites
复制标题

基于深度神经网络和生物子词的蛋白质S-硫基化位点检测

DOI:
10.1093/bib/bbaa128
复制
发表时间:
2021-05-01
影响因子:
9.5
通讯作者:
Nguyen Quoc Khanh Le
Nguyen Quoc Khanh Le
中科院分区:
生物学2区
文献类型:
--
作者:
Duyen Thi Do;Thanh Quynh Trang Le;Nguyen Quoc Khanh Le

文献摘要

被引文献

相似文献

蛋白质S-亚磺酰化是一种重要的翻译后修饰(PTMs),其羟基与半胱氨酸的巯基共价结合。最近的研究表明,这种修饰在信号转导、转录调控和细胞凋亡中起着重要的作用。到目前为止,蛋白质中次磺酸的动态仍然不清楚,因为它的短暂性质。因此,确定S-亚磺酰化位点可能是破译其神秘结构和功能的关键,这在细胞生物学和疾病中非常重要。然而,由于缺乏有效的方法,这一领域的科学家往往局限于一些耗时且不具成本效益的湿实验室技术。因此,这促使我们开发一种仅根据蛋白质序列信息检测S-亚硫烷基化位点的计算机模型。在这项研究中,蛋白质序列作为自然语言句子,包括生物子词。因此,深度神经网络被用于执行分类。独立数据集内的性能统计包括敏感性、特异性、准确性、马修斯相关系数和曲线下面积率分别达到85.71%、69.47%、77.09%、0.5554和0.833。我们的研究结果表明,与基准数据集上的其他知名工具相比,所提出的方法(fastSulf-DNN)在预测S-亚磺酰化位点方面取得了优异的性能。
Protein S-sulfenylation is one kind of crucial post-translational modifications (PTMs) in which the hydroxyl group covalently binds to the thiol of cysteine. Some recent studies have shown that this modification plays an important role in signaling transduction, transcriptional regulation and apoptosis. To date, the dynamic of sulfenic acids in proteins remains unclear because of its fleeting nature. Identifying S-sulfenylation sites, therefore, could be the key to decipher its mysterious structures and functions, which are important in cell biology and diseases. However, due to the lack of effective methods, scientists in this field tend to be limited in merely a handful of some wet lab techniques that are time-consuming and not cost-effective. Thus, this motivated us to develop an in silico model for detecting S-sulfenylation sites only from protein sequence information. In this study, protein sequences served as natural language sentences comprising biological subwords. The deep neural network was consequentially employed to perform classification. The performance statistics within the independent dataset including sensitivity, specificity, accuracy, Matthews correlation coefficient and area under the curve rates achieved 85.71%, 69.47%, 77.09%, 0.5554 and 0.833, respectively. Our results suggested that the proposed method (fastSulf-DNN) achieved excellent performance in predicting S-sulfenylation sites compared to other well-known tools on a benchmark dataset.