Filtering high-throughput protein-protein interaction data using a combination of genomic features.

Filtering high-throughput protein-protein interaction data using a combination of genomic features.
复制标题

DOI:
10.1186/1471-2105-6-100
复制
发表时间:
2005-04-18
期刊:
影响因子:
3
通讯作者:
Nakamura H
Nakamura H
中科院分区:
生物学4区
文献类型:
--
作者:
Patil A;Nakamura H

文献摘要

参考文献

被引文献

相似文献

蛋白质-蛋白质相互作用的数据用于分子网络的创建或预测,通常是从大规模或高通量的实验。这个实验数据很容易包含大量的虚假相互作用。因此,在预测研究中使用它们之前,需要验证相互作用并过滤掉不正确的数据。在这项研究中,我们使用3个基因组特征的组合-结构上已知的相互作用Pfam结构域,基因本体论注释和序列同源性-作为一种手段来分配可靠性的蛋白质-蛋白质相互作用在酿酒酵母高通量实验确定。使用贝叶斯网络的方法,我们表明,蛋白质-蛋白质相互作用的高通量数据支持的一个或多个基因组特征具有较高的似然比,因此更有可能是真实的相互作用。该方法具有较高的敏感性(90%)和良好的特异性(63%)。我们发现,56%的相互作用,从高通量实验在酿酒酵母具有很高的可靠性。我们使用该方法估计的真实相互作用的数量在高通量的蛋白质相互作用的数据集在秀丽隐杆线虫,果蝇和智人分别为27%,18%和68%。我们的结果可供搜索和下载。包括序列、结构和注释信息的基因组特征的组合是大的和有噪声的高通量数据集中真实相互作用的良好预测器。该方法具有非常高的灵敏度和良好的特异性,并可用于分配一个似然比,对应于可靠性,每个相互作用。
Protein-protein interaction data used in the creation or prediction of molecular networks is usually obtained from large scale or high-throughput experiments. This experimental data is liable to contain a large number of spurious interactions. Hence, there is a need to validate the interactions and filter out the incorrect data before using them in prediction studies. In this study, we use a combination of 3 genomic features – structurally known interacting Pfam domains, Gene Ontology annotations and sequence homology – as a means to assign reliability to the protein-protein interactions in Saccharomyces cerevisiae determined by high-throughput experiments. Using Bayesian network approaches, we show that protein-protein interactions from high-throughput data supported by one or more genomic features have a higher likelihood ratio and hence are more likely to be real interactions. Our method has a high sensitivity (90%) and good specificity (63%). We show that 56% of the interactions from high-throughput experiments in Saccharomyces cerevisiae have high reliability. We use the method to estimate the number of true interactions in the high-throughput protein-protein interaction data sets in Caenorhabditis elegans, Drosophila melanogaster and Homo sapiens to be 27%, 18% and 68% respectively. Our results are available for searching and downloading at . A combination of genomic features that include sequence, structure and annotation information is a good predictor of true interactions in large and noisy high-throughput data sets. The method has a very high sensitivity and good specificity and can be used to assign a likelihood ratio, corresponding to the reliability, to each interaction.
DOI: 10.1038/415141a
发表时间: 2002-01-10
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Bösche, M;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1101/gr.2203804
发表时间: 2004-06-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Asthana, S;King, OD;Roth, FP
通讯作者: Roth, FP
DOI: 10.1093/nar/30.1.31
发表时间: 2002-01-01
影响因子: 14.9
作者:
Mewes, HW;Frishman, D;Weil, B
通讯作者: Weil, B
DOI: 10.1101/gr.2122004
发表时间: 2004-07-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Lehner, B;Sanderson, CM
通讯作者: Sanderson, CM
DOI: 10.2142/biophysics.1.21
发表时间: 2005-01-01
期刊: Biophysics (Nagoya-shi, Japan)
影响因子: --
作者:
Patil, Ashwini;Nakamura, Haruki
通讯作者: Nakamura, Haruki