Data integration for plant genomics-exemplars from the integration of Arabidopsis thaliana databases

Data integration for plant genomics-exemplars from the integration of Arabidopsis thaliana databases
复制标题

DOI:
10.1093/bib/bbp047
复制
发表时间:
2009-11-01
影响因子:
9.5
通讯作者:
Rawlings, Christopher John
Rawlings, Christopher John
中科院分区:
生物学2区
文献类型:
--
作者:
Lysenko, Atem;Hindle, Matthew Morritt;Rawlings, Christopher John

文献摘要

被引文献

相似文献

植物科学问题的系统方法的发展需要整合现有的信息资源。然而,目前可用的信息往往是不完整的,分散在许多来源和数据的语法和语义异构性是一个挑战的集成。在这篇文章中,我们讨论了数据集成的策略,我们使用基于图形的集成方法(Ondex)来说明这些挑战中的一些参考两个例子的问题有关的整合(i)代谢途径和(ii)蛋白质相互作用的数据拟南芥。我们量化的重叠程度为三个常用的途径和蛋白质相互作用的信息来源。对于途径,我们发现AraCyc数据库包含最广泛的酶反应覆盖范围,对于蛋白质相互作用,我们发现IntAct数据库为综合数据集提供了最大的独特贡献。然而,对于这两个例子,我们观察到所有三个来源的共同数据量相对较小。综合网络的分析和视觉探索被用来确定一些与这些数据集的解释有关的实际问题。我们证明了这些方法的效用,从一个单独的微阵列实验的共表达基因组的分析,在通路信息的背景下,并与一个集成的蛋白质相互作用网络的共表达数据的组合。
The development of a systems based approach to problems in plant sciences requires integration of existing information resources. However, the available information is currently often incomplete and dispersed across many sources and the syntactic and semantic heterogeneity of the data is a challenge for integration. In this article, we discuss strategies for data integration and we use a graph based integration method (Ondex) to illustrate some of these challenges with reference to two example problems concerning integration of (i) metabolic pathway and (ii) protein interaction data for Arabidopsis thaliana. We quantify the degree of overlap for three commonly used pathway and protein interaction information sources. For pathways, we find that the AraCyc database contains the widest coverage of enzyme reactions and for protein interactions we find that the IntAct database provides the largest unique contribution to the integrated dataset. For both examples, however, we observe a relatively small amount of data common to all three sources. Analysis and visual exploration of the integrated networks was used to identify a number of practical issues relating to the interpretation of these datasets. We demonstrate the utility of these approaches to the analysis of groups of coexpressed genes from an individual microarray experiment, in the context of pathway information and for the combination of coexpression data with an integrated protein interaction network.