Comparing functional annotation analyses with Catmap -: art. no. 193

Comparing functional annotation analyses with Catmap -: art. no. 193
复制标题

DOI:
10.1186/1471-2105-5-193
复制
发表时间:
2004-12-09
期刊:
影响因子:
3
通讯作者:
Krogh, M
Krogh, M
中科院分区:
生物学4区
文献类型:
--
作者:
Breslin, T;Edén, P;Krogh, M

文献摘要

被引文献

相似文献

背景:来自微阵列实验的排序基因列表通常通过给预定义的基因类别赋予重要性来进行分析,例如。 g.,基于功能注释。执行此类分析的工具通常仅限于基于排名列表中的截止值的类别分数和基于随机基因排列作为零假设的显着性计算。结果:我们分析了三个公开可用的数据集,其中每个样本被分为两类,基因根据其与类标签的相关性进行排名。我们开发了一个程序 Catmap(可在 http://bioinfo.thep.lu.se/Catmap 下载),用于比较基因类别分析中的不同分数和零假设,使用基因本体注释进行类别定义。当使用基于截止值的分数时,结果在很大程度上取决于截止值的选择,从而在分析中引入任意性。分别比较使用随机基因排列和随机样本排列的结果,我们发现类别的指定显着性很大程度上取决于原假设的选择。与样本标签排列相比,对于具有许多共表达基因的大类别,基因排列给出的 p 值要小得多。结论:在排序基因列表的基因类别分析中,最好采用与截止无关的分数。原假设的选择非常重要;随机基因排列不能很好地近似样本标签排列。
Background: Ranked gene lists from microarray experiments are usually analysed by assigning significance to predefined gene categories, e. g., based on functional annotations. Tools performing such analyses are often restricted to a category score based on a cutoff in the ranked list and a significance calculation based on random gene permutations as null hypothesis.Results: We analysed three publicly available data sets, in each of which samples were divided in two classes and genes ranked according to their correlation to class labels. We developed a program, Catmap ( available for download at http://bioinfo.thep.lu.se/Catmap),to compare different scores and null hypotheses in gene category analysis, using Gene Ontology annotations for category definition. When a cutoff-based score was used, results depended strongly on the choice of cutoff, introducing an arbitrariness in the analysis. Comparing results using random gene permutations and random sample permutations, respectively, we found that the assigned significance of a category depended strongly on the choice of null hypothesis. Compared to sample label permutations, gene permutations gave much smaller p-values for large categories with many coexpressed genes.Conclusions: In gene category analyses of ranked gene lists, a cutoff independent score is preferable. The choice of null hypothesis is very important; random gene permutations does not work well as an approximation to sample label permutations.