Genome-wide matching of genes to cellular roles using guilt-by-association models derived from single sample analysis.

Genome-wide matching of genes to cellular roles using guilt-by-association models derived from single sample analysis.
复制标题

DOI:
10.1186/1756-0500-5-370
复制
发表时间:
2012-07-23
期刊:
影响因子:
1.8
通讯作者:
Furge KA
Furge KA
中科院分区:
其他
文献类型:
--
作者:
Klomp JA;Furge KA

文献摘要

被引文献

相似文献

确定每个基因产物的细胞或生理功能的高通量方法对于理解尚未被分子或遗传学方法广泛表征的基因的作用是有用的。推断基因功能的一种方法是“内疚-联想”,在这种方法中,特征不佳的基因的表达模式与特征较好的基因的表达方式是协同变化的。描述不充分的基因的功能是从描述良好的基因的已知功能(S)推断出来的。例如,与转录产物共表达的基因在细胞周期、发育、环境压力和肿瘤发生过程中都会发生变化,这些基因与这些过程有关。在研究几个特征不佳的基因的表达特征时,我们注意到,我们可以通过将单个基因的表达变化与单个样本的基因集丰富分数相关联,将每个基因与细胞表型联系起来。我们使用一个中等大小的基因表达数据集(EXPO)和一个基因表达表型概要(MSigDBv3.0)来评估该方法的有效性。我们发现,与线粒体和溶酶体基因集丰富相关最好的转录本大多与这些过程有关(分别为89/100和44/50)。互惠评估,根据丰富与单个基因表达的相关性对基因集进行排序,也反映了生物医学文献中突出基因的已知关联(16/19)。在对模型进行评估时,我们还发现4%的基因组编码与小分子和小肽信号转导基因集相关的蛋白质,这意味着大量的基因参与了内部和外部环境感知。我们的结果表明,这种方法对于推断不同基因集的功能是有用的。这种方法反映了其他人用来将单个基因与确定的基因表达变化联系起来的生物学实验方法。此外,该方法不仅可以用于发现与细胞过程相关的基因,还可以从与给定基因相关的提要中发现有意义的表达表型。这种方法的有效性、多功能性和广泛性使其能够在各种情况下和各种下游分析中应用。
High-throughput methods that ascribe a cellular or physiological function for each gene product are useful to understand the roles of genes that have not been extensively characterized by molecular or genetic approaches. One method to infer gene function is "guilt-by-association", in which the expression pattern of a poorly characterized gene is shown to co-vary with the expression of better-characterized genes. The function of the poorly characterized gene is inferred from the known function(s) of the well-described genes. For example, genes co-expressed with transcripts that vary during the cell cycle, development, environmental stresses, and with oncogenesis have been implicated in those processes. While examining the expression characteristics of several poorly characterized genes, we noted that we could associate each of the genes with a cellular phenotype by correlating individual gene expression changes with gene set enrichment scores from individual samples. We evaluated the effectiveness of this approach using a modest sized gene expression data set (expO) and a compendium of gene expression phenotypes (MSigDBv3.0). We found the transcripts that correlated best with enrichment in mitochondrial and lysosomal gene sets were mostly related to those processes (89/100 and 44/50, respectively). The reciprocal evaluation, ranking gene sets according to correlation of enrichment with an individual gene’s expression, also reflected known associations for prominent genes in the biomedical literature (16/19). In evaluating the model, we also found that 4% of the genome encodes proteins that are associated with small molecule and small peptide signal transduction gene sets, implicating a large number of genes in both internal and external environmental sensing. Our results show that this approach is useful to infer functions of disparate sets of genes. This method mirrors the biological experimental approaches used by others to associate individual genes with defined gene expression changes. Moreover, the approach can be used beyond discovering genes related to a cellular process to discover meaningful expression phenotypes from a compendium that are associated with a given gene. The effectiveness, versatility, and breadth of this approach make possible its application in a variety of contexts and with a variety of downstream analyses.