Accounting for redundancy when integrating gene interaction databases.

Accounting for redundancy when integrating gene interaction databases.
复制标题

DOI:
10.1371/journal.pone.0007492
复制
发表时间:
2009-10-22
期刊:
影响因子:
3.7
通讯作者:
Beyer A
Beyer A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Elefsinioti A;Ackermann M;Beyer A

文献摘要

参考文献

被引文献

相似文献

在过去的几年中,基因相互作用网络越来越多地被用于生物测量的评估和解释。了解未知蛋白质的相互作用伙伴可以让科学家了解遗传产物之间的复杂关系,有助于揭示未知的生物功能和途径,并更详细地了解生物体的复杂性。在所有相关条件下测量所有蛋白质相互作用几乎是不可能的。因此,需要整合不同数据集的计算方法来预测基因相互作用。然而,当整合不同的来源时,必须考虑到信息的某些部分可能是冗余的,这可能导致高估相互作用的真实可能性。我们的方法集成了来自三个不同的数据库(Bioverse,HiMAP和STRING)预测人类基因相互作用的信息。采用贝叶斯方法,以便在一个共同的量化尺度上整合不同的数据源。贝叶斯集成的一个重要假设是输入数据(特征)的独立性。我们的研究表明,条件依赖不能被忽略时,结合基因相互作用数据库,依赖于部分重叠的输入数据。此外,我们还展示了如何检测数据库之间的相关性结构,并提出了一个线性模型来纠正这种偏差。将结果与两个独立的参考数据集进行基准测试表明,集成模型的性能优于单个数据集。我们的方法提供了一个直观的策略,加权不同的功能,同时考虑他们的条件依赖性。
During the last years gene interaction networks are increasingly being used for the assessment and interpretation of biological measurements. Knowledge of the interaction partners of an unknown protein allows scientists to understand the complex relationships between genetic products, helps to reveal unknown biological functions and pathways, and get a more detailed picture of an organism's complexity. Being able to measure all protein interactions under all relevant conditions is virtually impossible. Hence, computational methods integrating different datasets for predicting gene interactions are needed. However, when integrating different sources one has to account for the fact that some parts of the information may be redundant, which may lead to an overestimation of the true likelihood of an interaction. Our method integrates information derived from three different databases (Bioverse, HiMAP and STRING) for predicting human gene interactions. A Bayesian approach was implemented in order to integrate the different data sources on a common quantitative scale. An important assumption of the Bayesian integration is independence of the input data (features). Our study shows that the conditional dependency cannot be ignored when combining gene interaction databases that rely on partially overlapping input data. In addition, we show how the correlation structure between the databases can be detected and we propose a linear model to correct for this bias. Benchmarking the results against two independent reference data sets shows that the integrated model outperforms the individual datasets. Our method provides an intuitive strategy for weighting the different features while accounting for their conditional dependencies.
DOI: 10.1093/nar/gkn760
发表时间: 2009-01
影响因子: 14.9
作者:
Jensen LJ;Kuhn M;Stark M;Chaffron S;Creevey C;Muller J;Doerks T;Julien P;Roth A;Simonovic M;Bork P;von Mering C
通讯作者: von Mering C
DOI: 10.1038/415141a
发表时间: 2002-01-10
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Bösche, M;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1093/nar/gkh036
发表时间: 2004-01-01
影响因子: 14.9
作者:
Harris, MA;Clark, J;White, R
通讯作者: White, R
DOI: 10.1038/47048
发表时间: 1999-11-04
期刊: NATURE
影响因子: 64.8
作者:
Marcotte, EM;Pellegrini, M;Eisenberg, D
通讯作者: Eisenberg, D
生物视频:增强蛋白质和蛋白质组结构,功能和上下文建模框架的增强。
DOI: 10.1093/nar/gki401
发表时间: 2005-07-01
影响因子: 14.9
作者:
McDermott, J;Guerquin, M;Frazier, Z;Chang, AN;Samudrala, R
通讯作者: Samudrala, R