Computation of significance scores of unweighted gene set enrichment analyses

Computation of significance scores of unweighted gene set enrichment analyses
复制标题

DOI:
10.1186/1471-2105-8-290
复制
发表时间:
2007-08-06
期刊:
影响因子:
3
通讯作者:
Lenhof, Hans-Peter
Lenhof, Hans-Peter
中科院分区:
生物学4区
文献类型:
--
作者:
Keller, Andreas;Backes, Christina;Lenhof, Hans-Peter

文献摘要

被引文献

相似文献

背景:基因集富集分析(Gene Set Enrichment Analysis,GSEA)是一种用于基因或蛋白质排序列表的统计评估的计算方法。最初GSEA是为了解释微阵列基因表达数据而开发的,但它可以应用于任何排序的基因列表。给定基因列表和任意生物类别,GSEA评估所考虑类别的基因是否随机分布或累积在列表的顶部或底部。通常情况下,显着性得分(p值)的GSEA计算的非参数排列检验,一个耗时的过程,只产生估计的p-values.Results:我们提出了一种新的动态规划算法计算精确的显着性值的未加权基因集富集分析。我们的算法避免了非参数排列检验的典型问题,即随机抽样过程导致不同运行中的结果不同。所提出的动态规划算法的另一个优点是它的运行时间和内存效率。为了测试我们的算法,我们不仅将其应用于模拟数据集,而且还评估了鳞状细胞肺癌组织和自体未受影响组织的表达谱。
Background: Gene Set Enrichment Analysis (GSEA) is a computational method for the statistical evaluation of sorted lists of genes or proteins. Originally GSEA was developed for interpreting microarray gene expression data, but it can be applied to any sorted list of genes. Given the gene list and an arbitrary biological category, GSEA evaluates whether the genes of the considered category are randomly distributed or accumulated on top or bottom of the list. Usually, significance scores (p-values) of GSEA are computed by nonparametric permutation tests, a time consuming procedure that yields only estimates of the p-values.Results: We present a novel dynamic programming algorithm for calculating exact significance values of unweighted Gene Set Enrichment Analyses. Our algorithm avoids typical problems of nonparametric permutation tests, as varying findings in different runs caused by the random sampling procedure. Another advantage of the presented dynamic programming algorithm is its runtime and memory efficiency. To test our algorithm, we applied it not only to simulated data sets, but additionally evaluated expression profiles of squamous cell lung cancer tissue and autologous unaffected tissue.