Analyzing yeast protein-protein interaction data obtained from different sources

Analyzing yeast protein-protein interaction data obtained from different sources
复制标题

DOI:
10.1038/nbt1002-991
复制
发表时间:
2002-10-01
影响因子:
46.9
通讯作者:
Hogue, CWV
Hogue, CWV
中科院分区:
工程技术1区
文献类型:
--
作者:
Bader, GD;Hogue, CWV

文献摘要

被引文献

相似文献

用于检测蛋白质相互作用的高通量方法,如质谱法和酵母双杂交测定法,继续产生大量的数据,可用于推断蛋白质的功能和调节。截至本文出版时,所有已发表的关于酿酒酵母的相互作用信息库是4,825种蛋白质之间的15,143种相互作用,幂律缩放支持20,000种特定蛋白质相互作用的估计。为了研究这些数据之间的偏差,重叠和互补性,我们对芽殖酵母中两个基于高通量质谱(HMS)的蛋白质相互作用数据集进行了分析,将它们相互比较并与其他相互作用数据集进行比较。我们的分析揭示了两个数据集共有的222种蛋白质之间的198种相互作用,其中许多反映了大的多蛋白复合物。它还表明,直接将诱饵蛋白与相关蛋白配对的“辐条”模型比连接所有蛋白质的“矩阵”模型准确约三倍。此外,我们确定了一个大的,以前未知的核仁复合物的148种蛋白质,其中包括39个未知功能的蛋白质。我们的研究结果表明,现有的大规模蛋白质相互作用数据集是不饱和的,整合许多不同的实验数据集产生一个更清晰的生物学观点比任何单一的方法。
High-throughput methods for detecting protein interactions, such as mass spectrometry and yeast two-hybrid assays, continue to produce vast amounts of data that may be exploited to infer protein function and regulation. As this article went to press, the pool of all published interaction information on Saccharomyces cerevisiae was 15,143 interactions among 4,825 proteins, and power-law scaling supports an estimate of 20,000 specific protein interactions. To investigate the biases, overlaps, and complementarities among these data, we have carried out an analysis of two high-throughput mass spectrometry (HMS)-based protein interaction data sets from budding yeast, comparing them to each other and to other interaction data sets. Our analysis reveals 198 interactions among 222 proteins common to both data sets, many of which reflect large multiprotein complexes. It also indicates that a "spoke" model that directly pairs bait proteins with associated proteins is roughly threefold more accurate than a "matrix" model that connects all proteins. In addition, we identify a large, previously unsuspected nucleolar complex of 148 proteins, including 39 proteins of unknown function. Our results indicate that existing large-scale protein interaction data sets are nonsaturating and that integrating many different experimental data sets yields a clearer biological view than any single method alone.