SPAR: a random forest-based predictor for self-interacting proteins with fine-grained domain information

SPAR: a random forest-based predictor for self-interacting proteins with fine-grained domain information
复制标题

SPAR:基于随机森林的具有细粒度域信息的自相互作用蛋白质预测器

DOI:
10.1007/s00726-016-2226-z
复制
发表时间:
2016-07-01
期刊:
影响因子:
3.5
通讯作者:
Song, Jiangning
Song, Jiangning
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, Xuhan;Yang, Shiping;Song, Jiangning

文献摘要

被引文献

相似文献

蛋白质自身相互作用,即同一基因表达的两个或多个相同蛋白质之间的相互作用,在细胞功能的调节中起着重要的作用。考虑到实验自相互作用鉴定的局限性,有必要设计特定的生物信息学工具来从蛋白质序列信息中预测自相互作用蛋白(SIP)。在这项研究中,我们提出了一种改进的SIP预测计算方法,称为SPAR(自交互蛋白质分析服务器)。首先,我们提出了一种改进的编码方案,称为关键残基替换(CRS),在该编码方案中考虑了细粒度的域-域相互作用信息。然后,利用随机森林算法对CRS的性能进行了评估,并与其他几种常用的基于序列的蛋白质-蛋白质相互作用预测编码方案进行了比较。通过在平衡训练数据集上的十次交叉验证测试,CRS表现最好,平均准确率达到72.01%。我们进一步将CRS与其他编码方案相结合,并使用最小冗余最大相关性(MRMR)特征选择方法来识别最重要的特征。我们选择的SPAR模型在非人类独立测试集上的平均准确率为92.09%(正负比约为1:11)。此外,我们还在一个独立的酵母测试集(正阴性比约为1:8)上对SPAR的性能进行了评估,获得了76.96%的平均准确率。结果表明,SPAR能够在跨物种应用中取得合理的性能。SPAR服务器可在http://systbio.cau.edu.cn/zzdlab/spar/上免费用于学术用途。
Protein self-interaction, i.e. the interaction between two or more identical proteins expressed by one gene, plays an important role in the regulation of cellular functions. Considering the limitations of experimental self-interaction identification, it is necessary to design specific bioinformatics tools for self-interacting protein (SIP) prediction from protein sequence information. In this study, we proposed an improved computational approach for SIP prediction, termed SPAR (Self-interacting Protein Analysis serveR). Firstly, we developed an improved encoding scheme named critical residues substitution (CRS), in which the fine-grained domain–domain interaction information was taken into account. Then, by employing the Random Forest algorithm, the performance of CRS was evaluated and compared with several other encoding schemes commonly used for sequence-based protein–protein interaction prediction. Through the tenfold cross-validation tests on a balanced training dataset, CRS performed the best, with the average accuracy up to 72.01 %. We further integrated CRS with other encoding schemes and identified the most important features using the mRMR (the minimum redundancy maximum relevance) feature selection method. Our SPAR model with selected features achieved an average accuracy of 92.09 % on the human-independent test set (the ratio of positives to negatives was about 1:11). Besides, we also evaluated the performance of SPAR on an independent yeast test set (the ratio of positives to negatives was about 1:8) and obtained an average accuracy of 76.96 %. The results demonstrate that SPAR is capable of achieving a reasonable performance in cross-species application. The SPAR server is freely available for academic use at http://systbio.cau.edu.cn/zzdlab/spar/ .