Mining functionally relevant gene sets for analyzing physiologically novel clinical expression data.

Mining functionally relevant gene sets for analyzing physiologically novel clinical expression data.
复制标题

挖掘功能相关的基因集以分析生理上新颖的临床表达数据。

DOI:
10.1142/9789814335058_0006
复制
发表时间:
2011
影响因子:
--
通讯作者:
Slonim,DonnaK
Slonim,DonnaK
中科院分区:
--
文献类型:
--
作者:
Turcan,Sevin;Vetter,DouglasE;Maron,JillL;Wei,Xintao;Slonim,DonnaK

文献摘要

被引文献

相似文献

基因集分析已成为提高转录组学研究灵敏度的标准方法。然而,结合基因集的分析方法需要与正在研究的潜在生理学相关的预定义基因集的可用性。对于新的生理问题,相关的基因集可能是不可用的,或现有的基因集数据库可能会偏向于只对相关的生物过程的最佳研究的结果。我们描述了一个成功的尝试,挖掘新的功能基因集的翻译项目的基础生理不一定是很好的特点,在现有的注释数据库。我们从公共表达数据库中选择有针对性的训练数据,并定义了选择双聚类作为候选基因集的新标准。许多发现的基因集显示很少或没有丰富的信息基因本体论术语或其他功能注释。然而,我们观察到这些基因集在新的临床测试数据集中显示出一致的差异表达,即使来自不同的物种、组织和疾病状态。我们证明了这种方法对人类代谢数据集的有效性,在那里我们发现了诊断糖尿病的新的,未表征的基因集,以及与神经元过程和人类发育相关的其他数据集。我们的研究结果表明,我们的方法可能是一种有效的方式来生成一个收集的基因集相关的新的临床应用的数据分析,现有的功能注释是相对不完整的。
Gene set analyses have become a standard approach for increasing the sensitivity of transcriptomic studies. However, analytical methods incorporating gene sets require the availability of pre-defined gene sets relevant to the underlying physiology being studied. For novel physiological problems, relevant gene sets may be unavailable or existing gene set databases may bias the results towards only the best-studied of the relevant biological processes. We describe a successful attempt to mine novel functional gene sets for translational projects where the underlying physiology is not necessarily well characterized in existing annotation databases. We choose targeted training data from public expression data repositories and define new criteria for selecting biclusters to serve as candidate gene sets. Many of the discovered gene sets show little or no enrichment for informative Gene Ontology terms or other functional annotation. However, we observe that such gene sets show coherent differential expression in new clinical test data sets, even if derived from different species, tissues, and disease states. We demonstrate the efficacy of this method on a human metabolic data set, where we discover novel, uncharacterized gene sets that are diagnostic of diabetes, and on additional data sets related to neuronal processes and human development. Our results suggest that our approach may be an efficient way to generate a collection of gene sets relevant to the analysis of data for novel clinical applications where existing functional annotation is relatively incomplete.