RANDOM-SET METHODS IDENTIFY DISTINCT ASPECTS OF THE ENRICHMENT SIGNAL IN GENE-SET ANALYSIS

RANDOM-SET METHODS IDENTIFY DISTINCT ASPECTS OF THE ENRICHMENT SIGNAL IN GENE-SET ANALYSIS
复制标题

DOI:
10.1214/07-aoas104
复制
发表时间:
2007-06-01
影响因子:
1.8
通讯作者:
Ahlquist, Paui
Ahlquist, Paui
中科院分区:
数学4区
文献类型:
--
作者:
Newton, Michael A.;Quintana, Fernando A.;Ahlquist, Paui

文献摘要

被引文献

相似文献

对于相对于细胞的两个或多个状态改变了表达水平的基因,可以不同程度地丰富预先指定的一组基因。了解由功能类别定义的基因集的丰富。例如基因本体论(GO)注释,对于分析微阵列表达数据中的生物信号是有价值的。衡量富集度的一种常见方法是根据功能类别中的成员资格和选定的显着变化基因列表对基因进行交叉分类。例如,在这个2x2表中,一个小的Fisher‘s精确检验P值就表示富集度。其他类别分析方法通过将类别级别的统计数据引用到与原始差异表达问题相关的排列分布来保留定量基因级别的分数并测量重要性。我们描述了一类随机集计分方法,该方法测量丰富信号的不同分量。这门课包括基于选定基因的费舍尔测试,也包括测试整个类别的平均基因水平证据。使用Affymetrix在鼻咽癌组织中表达的数据,并从理论上使用差异表达的定位模型,对平均和选择方法进行了实证比较。我们发现,在浓缩问题的状态空间中,每种方法都有各自的优势领域,而且两种方法在实际应用中都有一定的优势。我们的分析还解决了与多类别推理相关的两个问题,即如果相同丰富的类别具有不同的大小,则不会以相同的概率检测到它们,以及由于共享基因,类别统计之间存在依赖关系。随机集浓缩计算不需要蒙特卡罗来实现。它们在R包ALIZ中提供。
A prespecified set of genes may be enriched, to varying degrees, for genes that have altered expression levels relative to two or more states of a cell. Knowing the enrichment of gene sets defined by functional categories. such as gene ontology (GO) annotations, is valuable for analyzing the biological signals in microarray expression data. A common approach to measuring enrichment is by cross-classifying genes according to membership in a functional category and membership oil a selected list of significantly altered genes. A small Fisher's exact test P-value, for example, in this 2 x 2 table is indicative of enrichment. Other category analysis methods retain the quantitative gene-level scores and measure significance by referring a category-level statistic to a permutation distribution associated with the original differential expression problem. We describe a class of random-set scoring methods that measure distinct components of the enrichment signal. The class includes Fisher's test based on selected genes and also tests that average gene-level evidence across the category. Averaging and selection methods are compared empirically using Affymetrix data on expression in nasopharyngeal cancer tissue, and theoretically using a location model of differential, expression. We find that each method has a domain of superiority in the state space of enrichment problems, and that both methods have benefits in practice. Our analysis also addresses two problems related to multiple-category inference, namely, that equally enriched categories are not detected with equal probability if they are of different sizes, and also that there is dependence among category statistics owing to shared genes. Random-set enrichment calculations do not require Monte Carlo for implementation. They are made available in the R package allez.