Ontology-Based Meta-Analysis of Global Collections of High-Throughput Public Data

Ontology-Based Meta-Analysis of Global Collections of High-Throughput Public Data
复制标题

DOI:
10.1371/journal.pone.0013066
复制
发表时间:
2010-09-29
期刊:
影响因子:
3.7
通讯作者:
Ronaghi, Mostafa
Ronaghi, Mostafa
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kupershmidt, Ilya;Su, Qiaojuan Jane;Ronaghi, Mostafa

文献摘要

被引文献

相似文献

背景资料:如果我们要了解疾病的发展和设计有效的新疗法,就必须调查支配生物系统的分子和遗传事件之间的相互联系。微阵列和下一代测序技术有可能提供这些信息。然而,充分利用这些方法需要在大量高度异质的基因组数据集上建立生物学联系。利用公共领域中日益庞大的基因组数据正在迅速成为当今研究界的主要挑战之一。方法/结果:我们开发了一种新的数据挖掘框架,使研究人员能够使用这种不断增长的公共高通量数据来调查任何一组基因或蛋白质。来自微阵列和其他基因组平台的数千个异质数据集的分子状态之间的连接性通过基于秩的富集统计、荟萃分析和生物医学本体的组合来确定。我们通过数据集复制和荟萃分析解决数据质量问题,并确保大多数发现是使用多条证据线得出的。作为一个例子,我们的战略和效用的这个框架,我们应用我们的数据挖掘方法来探索生物学的棕色脂肪的背景下,数以千计的公开的基因表达datasets.Conclusions:我们的工作提出了一个实用的策略,组织,挖掘和相关的全球收集的大规模基因组数据,探索正常和疾病的生物学。使用无假设的方法,我们展示了如何在非常大的基因组数据集合中进行数据驱动的分析,以揭示新的发现和证据来支持现有的假设。
Background: The investigation of the interconnections between the molecular and genetic events that govern biological systems is essential if we are to understand the development of disease and design effective novel treatments. Microarray and next-generation sequencing technologies have the potential to provide this information. However, taking full advantage of these approaches requires that biological connections be made across large quantities of highly heterogeneous genomic datasets. Leveraging the increasingly huge quantities of genomic data in the public domain is fast becoming one of the key challenges in the research community today.Methodology/Results: We have developed a novel data mining framework that enables researchers to use this growing collection of public high-throughput data to investigate any set of genes or proteins. The connectivity between molecular states across thousands of heterogeneous datasets from microarrays and other genomic platforms is determined through a combination of rank-based enrichment statistics, meta-analyses, and biomedical ontologies. We address data quality concerns through dataset replication and meta-analysis and ensure that the majority of the findings are derived using multiple lines of evidence. As an example of our strategy and the utility of this framework, we apply our data mining approach to explore the biology of brown fat within the context of the thousands of publicly available gene expression datasets.Conclusions: Our work presents a practical strategy for organizing, mining, and correlating global collections of large-scale genomic data to explore normal and disease biology. Using a hypothesis-free approach, we demonstrate how a data-driven analysis across very large collections of genomic data can reveal novel discoveries and evidence to support existing hypothesis.