Integration of genomic datasets to predict protein complexes in yeast.

Integration of genomic datasets to predict protein complexes in yeast.
复制标题

DOI:
10.1023/a:1020495201615
复制
发表时间:
2002-01-01
期刊:
Journal of Structural and Functional Genomics
影响因子:
--
通讯作者:
Gerstein, Mark
Gerstein, Mark
中科院分区:
其他
文献类型:
--
作者:
Jansen, Ronald;Lan, Ning;Gerstein, Mark

文献摘要

被引文献

相似文献

功能基因组学的最终目标是确定生物体基因组中所有基因的功能。在过去几十年的研究中,已经积累和汇总了大量关于基因生物学作用的信息,这些信息既来自详细说明单个基因和蛋白质作用的传统实验,也来自旨在在基因组规模上表征基因功能的新实验策略。很明显,功能基因组学的目标只能通过整合来自这些不同实验的信息和数据源来实现。因此,不同数据的整合是生物信息学的一个重要挑战。不同数据源的整合通常有助于揭示基因之间不明显的关系,但还有两个进一步的好处。首先,只要来自多个独立来源的信息一致,就应该更加有效和可靠。其次,通过查看多个来源的联合,可以覆盖基因组的更大部分。这对于整合来自多个单基因或蛋白质实验的结果是显而易见的,但对于来自全基因组实验的许多结果也是必要的,因为它们通常局限于基因组的某些(尽管相当大)子集。在本文中,我们探讨了这样的数据集成过程的一个例子。我们专注于预测的成员在蛋白质复合物的个别基因。为此,我们招募了六个不同的数据源,包括表达谱,相互作用数据,必要性和定位信息。这些数据源中的每一个都单独包含一些关于蛋白质复合物的弱预测信息,但我们展示了如何通过组合所有这些数据源来改进这种预测。
The ultimate goal of functional genomics is to define the function of all the genes in the genome of an organism. A large body of information of the biological roles of genes has been accumulated and aggregated in the past decades of research, both from traditional experiments detailing the role of individual genes and proteins, and from newer experimental strategies that aim to characterize gene function on a genomic scale. It is clear that the goal of functional genomics can only be achieved by integrating information and data sources from the variety of these different experiments. Integration of different data is thus an important challenge for bioinformatics. The integration of different data sources often helps to uncover non-obvious relationships between genes, but there are also two further benefits. First, it is likely that whenever information from multiple independent sources agrees, it should be more valid and reliable. Secondly, by looking at the union of multiple sources, one can cover larger parts of the genome. This is obvious for integrating results from multiple single gene or protein experiments, but also necessary for many of the results from genome-wide experiments since they are often confined to certain (although sizable) subsets of the genome. In this paper, we explore an example of such a data integration procedure. We focus on the prediction of membership in protein complexes for individual genes. For this, we recruit six different data sources that include expression profiles, interaction data, essentiality and localization information. Each of these data sources individually contains some weakly predictive information with respect to protein complexes, but we show how this prediction can be improved by combining all of them.