Efficiently Predicting Hot Spots in PPIs by Combining Random Forest and Synthetic Minority Over-Sampling Technique

Efficiently Predicting Hot Spots in PPIs by Combining Random Forest and Synthetic Minority Over-Sampling Technique
复制标题

结合随机森林和合成少数过采样技术有效预测 PPI 中的热点

DOI:
10.1109/tcbb.2018.2871674
复制
发表时间:
2019-05
期刊:
IEEE/ACM Transactions on Computational Biology and Bioinformatics
影响因子:
--
通讯作者:
Xu Xin
Xu Xin
中科院分区:
其他
文献类型:
--
作者:
Zhang Xiaolong;Lin Xiaoli;Zhao Jiafu;Huang Qianqian;Xu Xin

文献摘要

参考文献

相似文献

热点残基在生物信息学中对发现新药、药物设计等方面发挥着重要作用。然而,目前的数据集主要由非热点组成,只有很小比例的热点。传统的热点预测方法面临着训练样本不均衡的问题。本文提出了一种结合随机森林分类和过采样策略的分类方法,以提高训练性能。使用具有过采样能力的策略来生成热点数据以平衡给定的训练集。然后调用随机森林分类来为该过采样训练集生成一组森林树。在过采样和训练过程之后,可以递归地计算最终的预测性能。该方法能够随机选择特征并构建一个鲁棒的随机森林,以避免过拟合训练集。三个数据集的实验结果表明,与现有的分类方法相比,热点预测的性能得到了显着改善。
Hot spot residues bring into play the vital function in bioinformatics to find new medications such as drug design. However, current datasets are predominately composed of non-hot spots with merely a tiny percentage of hot spots. Conventional hot spots prediction methods may face great challenges towards the problem of imbalance training samples. This paper presents a classification method combining with random forest classification and oversampling strategy to improve the training performance. A strategy with an oversampling ability is used to generate hot spots data to balance the given training set. Random forest classification is then invoked to generate a set of forest trees for this oversampled training set. The final prediction performance can be computed recursively after the oversampling and training process. This proposed method is capable of randomly selecting features and constructing a robust random forest to avoid overfitting the training set. Experimental results from three data sets indicate that the performance of hot spots prediction has been significantly improved compared with existing classification methods.
DOI: 10.1016/s0022-2836(02)00442-4
发表时间: 2002-07-05
影响因子: 5.6
作者:
Guerois, R;Nielsen, JE;Serrano, L
通讯作者: Serrano, L
使用局部调整禁忌搜索算法预测蛋白质结构
DOI: 10.1186/1471-2105-15-s15-s1
发表时间: 2014
期刊: BMC bioinformatics
影响因子: 3
作者:
Lin X;Zhang X;Zhou F
通讯作者: Zhou F
DOI: 10.1371/journal.pone.0016774
发表时间: 2011-02-28
期刊: PloS one
影响因子: 3.7
作者:
Lise S;Buchan D;Pontil M;Jones DT
通讯作者: Jones DT
DOI: 10.1109/tcbb.2013.10
发表时间: 2013-03-01
影响因子: 4.5
作者:
Huang, De-Shuang;Yu, Hong-Jie
通讯作者: Yu, Hong-Jie
DOI: 10.1007/978-1-4419-9326-7_5
发表时间: 2012-01-01
期刊: ENSEMBLE MACHINE LEARNING: METHODS AND APPLICATIONS
影响因子: --
作者:
Cutler, Adele;Cutler, D. Richard;Stevens, John R.
通讯作者: Stevens, John R.