iPPBS-Opt: A Sequence-Based Ensemble Classifier for Identifying Protein-Protein Binding Sites by Optimizing Imbalanced Training Datasets.

iPPBS-Opt: A Sequence-Based Ensemble Classifier for Identifying Protein-Protein Binding Sites by Optimizing Imbalanced Training Datasets.
复制标题

DOI:
10.3390/molecules21010095
复制
发表时间:
2016-01-19
期刊:
Molecules (Basel, Switzerland)
影响因子:
--
通讯作者:
Chou KC
Chou KC
中科院分区:
其他
文献类型:
--
作者:
Jia J;Liu Z;Xiao X;Liu B;Chou KC

文献摘要

参考文献

被引文献

相似文献

蛋白质-蛋白质相互作用及其结合位点的知识对于深入了解活细胞中的网络是必不可少的。随着后基因组时代产生的大量蛋白质序列,开发基于序列信息及时识别蛋白质-蛋白质结合位点(PPBS)的计算方法是至关重要的,因为通过这种方式获得的信息可以用于生物医学研究和药物开发。为了解决这一挑战,我们提出了一种新的预测器,称为iPPBS-Opt,其中我们使用了:(1)K最近邻清洗(KNNC)和插入假设训练样本(IHTS)处理来优化训练数据集;(2)集成投票方法来选择最相关的特征;(3)静态小波变换来制定统计样本。交叉验证测试表明,新的预测器是非常有前途的,这意味着上述做法确实是非常有效的。特别是,使用小波表示蛋白质/肽序列的方法可能是抓住问题本质的关键,这与蛋白质的许多重要生物功能可以用其低频内部运动来阐明的发现完全一致。为了最大限度地方便大多数实验科学家,我们提供了一个分步指南,介绍如何使用预测器的Web服务器()来获得所需的结果,而无需通过复杂的数学方程。
Knowledge of protein-protein interactions and their binding sites is indispensable for in-depth understanding of the networks in living cells. With the avalanche of protein sequences generated in the postgenomic age, it is critical to develop computational methods for identifying in a timely fashion the protein-protein binding sites (PPBSs) based on the sequence information alone because the information obtained by this way can be used for both biomedical research and drug development. To address such a challenge, we have proposed a new predictor, called iPPBS-Opt, in which we have used: (1) the K-Nearest Neighbors Cleaning (KNNC) and Inserting Hypothetical Training Samples (IHTS) treatments to optimize the training dataset; (2) the ensemble voting approach to select the most relevant features; and (3) the stationary wavelet transform to formulate the statistical samples. Cross-validation tests by targeting the experiment-confirmed results have demonstrated that the new predictor is very promising, implying that the aforementioned practices are indeed very effective. Particularly, the approach of using the wavelets to express protein/peptide sequences might be the key in grasping the problem’s essence, fully consistent with the findings that many important biological functions of proteins can be elucidated with their low-frequency internal motions. To maximize the convenience of most experimental scientists, we have provided a step-by-step guide on how to use the predictor’s web server () to get the desired results without the need to go through the complicated mathematical equations involved.
DOI: 10.1042/bj1870829
发表时间: 1980-01-01
影响因子: 4.1
作者:
CHOU, KC;FORSEN, S
通讯作者: FORSEN, S
DOI: 10.1093/bioinformatics/bth466
发表时间: 2005-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chou, KC
通讯作者: Chou, KC
DOI: 10.1016/j.ab.2015.08.021
发表时间: 2015-12-01
影响因子: 2.9
作者:
Chen, Wei;Feng, Pengmian;Chou, Kuo-Chen
通讯作者: Chou, Kuo-Chen
DOI: 10.1016/j.ab.2014.04.001
发表时间: 2014-07-01
影响因子: 2.9
作者:
Chen, Wei;Lei, Tian-Yu;Chou, Kuo-Chen
通讯作者: Chou, Kuo-Chen
DOI: 10.1002/bip.360340114
发表时间: 1994-01-01
期刊: BIOPOLYMERS
影响因子: 2.9
作者:
CHOU, KC;ZHANG, CT;MAGGIORA, GM
通讯作者: MAGGIORA, GM