Predicting protein-protein interactions in unbalanced data using the primary structure of proteins.

Predicting protein-protein interactions in unbalanced data using the primary structure of proteins.
复制标题

DOI:
10.1186/1471-2105-11-167
复制
发表时间:
2010-04-02
期刊:
影响因子:
3
通讯作者:
Chang DT
Chang DT
中科院分区:
生物学4区
文献类型:
--
作者:
Yu CY;Chou LC;Chang DT

文献摘要

参考文献

被引文献

相似文献

阐明蛋白质-蛋白质相互作用(PPI)对于构建蛋白质相互作用网络和促进我们理解生物系统的一般原理至关重要。先前的研究表明,相互作用的蛋白质对可以通过其一级结构来预测。这些方法中的大多数在包含相同数量的相互作用和非相互作用蛋白质对的数据集上取得了令人满意的性能。然而,这个比率本质上是高度不平衡的,并且这些技术尚未针对现实数据集中大量非交互对的影响进行全面评估。此外,由于高度不平衡的分布通常会导致数据集很大,因此在处理此类具有挑战性的任务时需要更有效的预测器。本研究提出了一种仅基于序列信息的PPI预测方法,该方法在三个方面做出了贡献。首先,我们提出了一种基于概率的机制,用于将蛋白质序列转换为特征向量。其次,所提出的预测器是采用高效的分类算法设计的,其中效率对于处理高度不平衡的数据集至关重要。第三,所提出的 PPI 预测器使用多个具有不同正负比(从 1:1 到 1:15)的不平衡数据集进行评估。该分析提供了确凿的证据,表明数据集不平衡的程度对于 PPI 预测因子很重要。处理数据不平衡是 PPI 预测中的一个关键问题,因为相互作用的蛋白质对比非相互作用的蛋白质对少得多。本文对这个问题进行了全面的研究,并开发了一种实用工具,仅使用蛋白质序列信息即可实现良好的预测性能和效率。
Elucidating protein-protein interactions (PPIs) is essential to constructing protein interaction networks and facilitating our understanding of the general principles of biological systems. Previous studies have revealed that interacting protein pairs can be predicted by their primary structure. Most of these approaches have achieved satisfactory performance on datasets comprising equal number of interacting and non-interacting protein pairs. However, this ratio is highly unbalanced in nature, and these techniques have not been comprehensively evaluated with respect to the effect of the large number of non-interacting pairs in realistic datasets. Moreover, since highly unbalanced distributions usually lead to large datasets, more efficient predictors are desired when handling such challenging tasks. This study presents a method for PPI prediction based only on sequence information, which contributes in three aspects. First, we propose a probability-based mechanism for transforming protein sequences into feature vectors. Second, the proposed predictor is designed with an efficient classification algorithm, where the efficiency is essential for handling highly unbalanced datasets. Third, the proposed PPI predictor is assessed with several unbalanced datasets with different positive-to-negative ratios (from 1:1 to 1:15). This analysis provides solid evidence that the degree of dataset imbalance is important to PPI predictors. Dealing with data imbalance is a key issue in PPI prediction since there are far fewer interacting protein pairs than non-interacting ones. This article provides a comprehensive study on this issue and develops a practical tool that achieves both good prediction performance and efficiency using only protein sequence information.
DOI: 10.1093/nar/gkj003
发表时间: 2006-01-01
影响因子: 14.9
作者:
Güldener U;Münsterkötter M;Oesterheld M;Pagel P;Ruepp A;Mewes HW;Stümpflen V
通讯作者: Stümpflen V
DOI: 10.1186/gb-2006-7-11-120
发表时间: 2006
期刊: Genome biology
影响因子: 12.3
作者:
Hart GT;Ramani AK;Marcotte EM
通讯作者: Marcotte EM
DOI: 10.1038/415141a
发表时间: 2002-01-10
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Bösche, M;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1093/bioinformatics/bth366
发表时间: 2004-11-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Huang, TW;Tien, AC;Huang, CYF
通讯作者: Huang, CYF
DOI: 10.1038/nature04532
发表时间: 2006-03-30
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Aloy, P;Superti-Furga, G
通讯作者: Superti-Furga, G