Improved Prediction of Protein-Protein Interaction Mapping on Homo Sapiens by Using Amino Acid Sequence Features in a Supervised Learning Framework

Improved Prediction of Protein-Protein Interaction Mapping on Homo Sapiens by Using Amino Acid Sequence Features in a Supervised Learning Framework
复制标题

DOI:
10.2174/0929866527666200610141258
复制
发表时间:
2021-01-01
影响因子:
1.6
通讯作者:
Mollah, Md Nurul Haque
Mollah, Md Nurul Haque
中科院分区:
生物学4区
文献类型:
--
作者:
Islam, Md Merajul;Alam, Md Jahangir;Mollah, Md Nurul Haque

文献摘要

被引文献

相似文献

背景:蛋白质-蛋白质相互作用(PPI)已成为控制许多生物过程(包括蛋白质功能、疾病发生率和治疗设计)的关键作用。然而,通过湿实验室实验鉴定 PPI 是一项具有挑战性的任务,因为它费力、耗时且昂贵。因此,在进行实验验证之前,现在重点关注 PPI 的计算预测,因为它同时省力、省时和成本最小化。 目的:本研究的目的是通过使用监督学习框架中的氨基酸序列特征,开发一种改进的计算方法,用于在智人上进行 PPI 预测图谱。方法:从 IntAct 分子相互作用数据库中收集经过实验验证的 91 个阳性 PPI 对人类蛋白质序列。然后我们构建了三个平衡数据集,正负 PPI 样本比例为 1:1、1:2 和 1:3。然后,我们将每个数据集划分为训练数据集 (80%) 和独立测试数据集 (20%)。同样,每个训练数据集被分为四个大小相等的互斥组,以便将每个组与独立的测试组互换,以执行 5 倍交叉验证 (CV)。然后,我们用每种比率情况训练候选七个分类器(NN、SVM、LR、NB、KNN、AB 和 RF),通过比较它们的性能得分来获得更好的 PPI 预测器。结果:基于 AAC 编码特征的正 PPI 和负 PPI 样本比例为 1:2 进行训练的基于随机森林 (RF) 的预测器通过产生准确度 (93.50%)、灵敏度 (95.0%) 和灵敏度 (95.0%) 的最高平均性能得分,提供了最准确的 PPI 预测。 5 倍交叉验证的 MCC (85.2%)、AUC (0.941) 和 pAUC (0.236)。与其他候选预测器和现有预测器进行比较时,它在独立测试数据集上还获得了准确率 (92.0%)、灵敏度 (94.0%)、MCC (83.6%)、AUC (0.922) 和 pAUC (0.207) 的最高平均性能分数。结论:最终结果预测强烈推荐基于 RF 的预测器是智人 PPI 映射的更好预测模型。
Background: Protein-Protein Interaction (PPI) has emerged as a key role in the control of many biological processes including protein function, disease incidence, and therapy design. However, the identification of PPI by wet lab experiment is a challenging task, since it is laborious, time consuming and expensive. Therefore, computational prediction of PPI is now given emphasis before going to the experimental validation, since it is simultaneously less laborious, time saver and cost minimizer.Objective: The objective of this study is to develop an improved computational method for PPI prediction mapping on Homo sapiens by using the amino acid sequence features in a supervised learning framework.Methods: The experimentally validated 91 positive-PPI pairs of human protein sequences were collected from IntAct Molecular Interaction Database. Then we constructed three balanced datasets with ratios 1:1, 1:2 and 1:3 of positive and negative PPI samples. Then we partitioned each dataset into training (80%) and independent test (20%) datasets. Again each training dataset was partitioned into four mutually exclusive groups of equal sizes for interchanging each group with independent test group to perfonn 5-fold cross validation (CV). Then we trained candidate seven classifiers (NN, SVM, LR, NB, KNN, AB and RF) with each ratio case to obtain the better PPI predictor by comparing their performance scores.Results: The random forest (RF) based predictor that was trained with 1:2 ratio of positive-PPI and negative-PPI samples based on AAC encoding features provided the most accurate PPI prediction by producing the highest average performance scores of accuracy (93.50%), sensitivity (95.0%), MCC (85.2%), AUC (0.941) and pAUC (0.236) with the 5-fold cross-validation. It also achieved the highest average performance scores of accuracy (92.0%), sensitivity (94.0%), MCC (83.6%), AUC (0.922) and pAUC (0.207) with the independent test datasets in a comparison of the other candidate and existing predictors.Conclusion: The final resultant prediction strongly recommend that the RF based predictor is a better prediction model of PPI mapping on Homo sapiens.