Generalized random set framework for functional enrichment analysis using primary genomics datasets

Generalized random set framework for functional enrichment analysis using primary genomics datasets
复制标题

DOI:
10.1093/bioinformatics/btq593
复制
发表时间:
2011-01-01
期刊:
影响因子:
5.8
通讯作者:
Medvedovic, Mario
Medvedovic, Mario
中科院分区:
生物学3区
文献类型:
--
作者:
Freudenberg, Johannes M.;Sivaganesan, Siva;Medvedovic, Mario

文献摘要

被引文献

相似文献

动机:使用原始基因组数据集进行功能富集分析是一种新兴方法,可以补充基于预定义的功能相关基因列表的功能富集现有方法。目前使用的方法取决于根据临时显着性截止值创建“显着”和“非显着”基因列表。这可能会导致统计功效的损失,并可能引入影响实验结果解释的偏差。结果:我们开发并验证了一种新的统计框架,即广义随机集(GRS)分析,用于比较两个数据集中的基因组特征,而无需进行基因分类。在我们的测试中,GRS 产生了正确的统计显着性度量,并且与当前在此环境中使用的其他方法相比,它的统计功效显着提高。我们还开发了一种用于识别驱动基因组图谱一致性的基因的程序,并证明了在此类分析中识别的基因的功能一致性的显着改善。
Motivation: Functional enrichment analysis using primary genomics datasets is an emerging approach to complement established methods for functional enrichment based on predefined lists of functionally related genes. Currently used methods depend on creating lists of 'significant' and 'non-significant' genes based on ad hoc significance cutoffs. This can lead to loss of statistical power and can introduce biases affecting the interpretation of experimental results.Results: We developed and validated a new statistical framework, generalized random set (GRS) analysis, for comparing the genomic signatures in two datasets without the need for gene categorization. In our tests, GRS produced correct measures of statistical significance, and it showed dramatic improvement in the statistical power over other methods currently used in this setting. We also developed a procedure for identifying genes driving the concordance of the genomics profiles and demonstrated a dramatic improvement in functional coherence of genes identified in such analysis.