Genome-wide inference of protein interaction sites: lessons from the yeast high-quality negative protein-protein interaction dataset.

Genome-wide inference of protein interaction sites: lessons from the yeast high-quality negative protein-protein interaction dataset.
复制标题

蛋白质相互作用位点的全基因组推断:来自酵母高质量阴性蛋白质 - 蛋白质相互作用数据集的经验教训。

DOI:
10.1093/nar/gkn016
复制
发表时间:
2008-04
影响因子:
14.9
通讯作者:
Lin, Kui
Lin, Kui
中科院分区:
生物学2区
文献类型:
--
作者:
Guo, Jie;Wu, Xiaomei;Zhang, Da-Yong;Lin, Kui

文献摘要

参考文献

被引文献

相似文献

蛋白质相互作用的高通量研究可能已经在实验和计算上产生了完全测序基因组中最全面的蛋白质-蛋白质相互作用数据集。它为我们提供了一个在蛋白质组尺度上发现潜在的蛋白质相互作用模式的机会。在这里,我们提出了一种在相互作用位点(通常为3-8个残基)上发现基序对的方法,这对于理解蛋白质功能和蛋白质工程和折叠实验的合理设计至关重要。挖掘一个金标准正(相互作用)数据集和一个金标准负(非相互作用)数据集来推断相互作用的基序对,这些基序对在正数据集中比在负数据集中明显被过度代表。对不同策略组合的4个阴性数据集进行评估,并将表现最佳的一个作为金标准阴性进行进一步分析。同时,为了评估我们的方法在检测潜在相互作用基序对方面的效率,我们比较了之前开发的其他方法,发现我们的方法达到了最高的预测精度。此外,许多未表征的基序对在其他物种中被发现具有实验证据。这项调查证明了高质量的负数据集对这种统计推断的性能的重要影响。
High-throughput studies of protein interactions may have produced, experimentally and computationally, the most comprehensive protein–protein interaction datasets in the completely sequenced genomes. It provides us an opportunity on a proteome scale, to discover the underlying protein interaction patterns. Here, we propose an approach to discovering motif pairs at interaction sites (often 3–8 residues) that are essential for understanding protein functions and helpful for the rational design of protein engineering and folding experiments. A gold standard positive (interacting) dataset and a gold standard negative (non-interacting) dataset were mined to infer the interacting motif pairs that are significantly overrepresented in the positive dataset compared to the negative dataset. Four negative datasets assembled by different strategies were evaluated and the one with the best performance was used as the gold standard negatives for further analysis. Meanwhile, to assess the efficiency of our method in detecting potential interacting motif pairs, other approaches developed previously were compared, and we found that our method achieved the highest prediction accuracy. In addition, many uncharacterized motif pairs of interest were found to be functional with experimental evidence in other species. This investigation demonstrates the important effects of a high-quality negative dataset on the performance of such statistical inference.
DOI: 10.1186/gb-2006-7-11-120
发表时间: 2006
期刊: Genome biology
影响因子: 12.3
作者:
Hart GT;Ramani AK;Marcotte EM
通讯作者: Marcotte EM
DOI: 10.1371/journal.pcbi.0020124
发表时间: 2006-09-29
影响因子: 4.3
作者:
Kim WK;Henschel A;Winter C;Schroeder M
通讯作者: Schroeder M
DOI: 10.1093/bioinformatics/btl020
发表时间: 2006-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, HQ;Li, JY;Wong, LS
通讯作者: Wong, LS
DOI: 10.1126/science.1105103
发表时间: 2005-02-04
期刊: SCIENCE
影响因子: 56.9
作者:
de Lichtenberg, U;Jensen, LJ;Bork, P
通讯作者: Bork, P
DOI: 10.1038/35075138
发表时间: 2001-05-03
期刊: NATURE
影响因子: 64.8
作者:
Jeong, H;Mason, SP;Oltvai, ZN
通讯作者: Oltvai, ZN