Revisiting the negative example sampling problem for predicting protein-protein interactions

Revisiting the negative example sampling problem for predicting protein-protein interactions
复制标题

DOI:
10.1093/bioinformatics/btr514
复制
发表时间:
2011-11-01
期刊:
影响因子:
5.8
通讯作者:
Marcotte, Edward M.
Marcotte, Edward M.
中科院分区:
生物学3区
文献类型:
--
作者:
Park, Yungki;Marcotte, Edward M.

文献摘要

被引文献

相似文献

动机:已经提出了许多计算方法,预测蛋白质-蛋白质相互作用(PPI)的基础上蛋白质序列的功能。由于潜在的非相互作用蛋白质对(负PPI)的数量在绝对值上和与相互作用蛋白质对(正PPI)的数量相比都非常高,因此计算预测方法依赖于负PPI的子集进行训练和验证。因此,需要出现的子集抽样为负PPIS.Results:我们澄清,有两个根本不同类型的子集抽样为负PPIs。一种是交叉验证测试的子集抽样,其中一个期望无偏子集,以便可以安全地假设用它们估计的预测性能推广到总体水平。另一种是用于训练的子集采样,其中人们希望最好地训练预测算法的子集,即使这些子集有偏差。我们发现,这两种根本不同类型的子集采样之间的混淆导致最近发表在《生物信息学》上的一项研究得出错误的结论,即基于蛋白质序列特征的预测算法在预测PPI方面几乎不比随机算法好。相反,蛋白质序列特征和相互作用蛋白质的“hubbiness”有助于有效预测PPI。我们提供指导,以适当使用随机与平衡抽样。
Motivation: A number of computational methods have been proposed that predict protein-protein interactions (PPIs) based on protein sequence features. Since the number of potential non-interacting protein pairs ( negative PPIs) is very high both in absolute terms and in comparison to that of interacting protein pairs ( positive PPIs), computational prediction methods rely upon subsets of negative PPIs for training and validation. Hence, the need arises for subset sampling for negative PPIs.Results: We clarify that there are two fundamentally different types of subset sampling for negative PPIs. One is subset sampling for cross-validated testing, where one desires unbiased subsets so that predictive performance estimated with them can be safely assumed to generalize to the population level. The other is subset sampling for training, where one desires the subsets that best train predictive algorithms, even if these subsets are biased. We show that confusion between these two fundamentally different types of subset sampling led one study recently published in Bioinformatics to the erroneous conclusion that predictive algorithms based on protein sequence features are hardly better than random in predicting PPIs. Rather, both protein sequence features and the 'hubbiness' of interacting proteins contribute to effective prediction of PPIs. We provide guidance for appropriate use of random versus balanced sampling.