UbiSitePred: A novel method for improving the accuracy of ubiquitination sites prediction by using LASSO to select the optimal Chou's pseudo components

UbiSitePred: A novel method for improving the accuracy of ubiquitination sites prediction by using LASSO to select the optimal Chou's pseudo components
复制标题

UbiSitePred:一种通过使用 LASSO 选择最佳 Chou 伪组件来提高泛素化位点预测准确性的新方法

DOI:
10.1016/j.chemolab.2018.11.012
复制
发表时间:
2019-01-15
影响因子:
3.9
通讯作者:
Ma, Qin
Ma, Qin
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cui, Xiaowen;Yu, Zhaomin;Ma, Qin

文献摘要

被引文献

相似文献

泛素化是蛋白质翻译后修饰的重要过程,在蛋白酶体降解、转录调控、DNA损伤修复等细胞生命活动中起着至关重要的作用。因此,识别泛素化位点是理解泛素化分子机制的关键步骤。然而,大量的泛素化位点的实验验证是耗时和昂贵的。为了缓解这些问题,需要一种计算方法来预测泛素化位点。提出了一种结合最小绝对收缩选择算子(LASSO)特征选择和支持向量机的泛素化位点预测新方法UbiSitePred。首先,利用二进制编码(BE)、伪氨基酸组成(PseAAC)、k间隔氨基酸对组成(CKSAAP)、位置特异性倾向矩阵(PSPM)提取序列特征信息,得到初始特征空间。其次,应用LASSO算法去除特征冗余信息,选择最优特征子集。最后,将最优特征子集输入支持向量机(SVM)进行泛素化位点预测。五重交叉验证结果表明,UbiSitePred模型的预测性能优于其他方法,Set1、Set2和Set3的AUC值分别为0.9998、0.8887和0.8481。值得注意的是,UbiSitePred的总体准确率分别为98.33%、81.12%和76.90%。实验结果表明,该方法明显上级现有的预测方法,为蛋白质翻译后修饰位点的预测提供了新的思路.源代码和所有数据集都可以在https://github.com/QUST-AIBBDRC/UbiSitePred/上找到。
Ubiquitination is an essential process in protein post-translational modification, which plays a crucial role in cell life activities, such as proteasomal degradation, transcriptional regulation, and DNA damage repair. Therefore, recognition of ubiquitination sites is a crucial step to understand the molecular mechanisms of ubiquitination. However, the experimental verification of numerous ubiquitination sites is time-consuming and costly. To alleviate these issues, a computational approach is needed to predict ubiquitination sites. This paper proposes a new method called UbiSitePred for predicting ubiquitination sites combined least absolute shrinkage and selection operator (LASSO) feature selection and support vector machine. First, we use binary encoding (BE), pseudo-amino acid composition (PseAAC), the composition of k-spaced amino acid pairs (CKSAAP), position-specific propensity matrices (PSPM) to extract the sequence feature information; thus, the initial feature space is obtained. Secondly, LASSO is applied to remove the feature redundancy information and selects the optimal feature subset. Finally, the optimal feature subset is input into the support vector machine (SVM) to predict the ubiquitination sites. Five-fold cross-validation shows that UbiSitePred model can achieve a better prediction performance compared with other methods, the AUC values for Set1, Set2, and Set3 are 0.9998, 0.8887, and 0.8481, respectively. Notably, the UbiSitePred has overall accuracy rates of 98.33%, 81.12%, and 76.90%, respectively. The results demonstrate that the proposed method is significantly superior to other state-of-the-art prediction methods and provide a new idea for the prediction of other post-translational modification sites of proteins. The source code and all datasets are available at https://github.com/QUST-AIBBDRC/UbiSitePred/.