Optimal sample size for multiple testing:: The case of gene expression microarrays

Optimal sample size for multiple testing:: The case of gene expression microarrays
复制标题

DOI:
10.1198/016214504000001646
复制
发表时间:
2004-12-01
影响因子:
3.7
通讯作者:
Rousseau, J
Rousseau, J
中科院分区:
数学1区
文献类型:
--
作者:
M端ller, P;Parmigiani, G;Rousseau, J

文献摘要

被引文献

相似文献

我们考虑多重比较问题的最佳样本量的选择。激励应用是在学习差异基因表达时要进行的微阵列实验数量的选择。然而,该方法在任何涉及大量假设检验的多重比较的应用中都是有效的。我们在此设置的上下文中讨论两个决策问题:样本量的选择和多重比较的决定。我们采用决策理论方法,使用损失函数,将发现尽可能多的差异表达基因的竞争目标结合起来,同时保持错误发现的数量可控。为了一致性,我们对两个决策使用相同的损失函数。多重比较问题的决策规则采用了最近文献中提出的控制后验期望错误发现率的规则的精确形式。对于样本量的选择,我们将期望效用参数与额外的敏感性分析结合起来,报告条件期望效用和对真实差分表达式的假设水平的调节。我们承认由此产生的诊断是一种促进解释和交流的统计力量。作为跨基因和阵列观察到的基因表达密度的抽样模型,我们使用了分层伽玛/伽玛模型的变体。但决策问题的讨论与选择的概率模型无关。该方法适用于任何模型,包括多重比较中零假设的正先验概率,并允许有效的边际和后验模拟,可能通过依赖马尔可夫链蒙特卡罗模拟。
We consider the choice of an optimal sample size for multiple-comparison problems. The motivating application is the choice of the number of microarray experiments to be carried out when learning about differential gene expression. However, the approach is valid in any application that involves multiple comparisons in a large number of hypothesis tests. We discuss two decision problems in the context of this setup: the. sample size selection and the decision about the multiple comparisons. We adopt a decision-theoretic approach, using loss functions that combine the competing goals of discovering as many differentially expressed genes as possible, while keeping the number of false discoveries manageable. For consistency, we use the same loss function for both decisions. The decision rule that emerges for the multiple-comparison problem takes the exact form of the rules proposed in the recent literature to control the posterior expected false-discovery rate. For the sample size selection, we combine the expected utility argument with an additional sensitivity analysis, reporting the conditional expected utilities and conditioning on assumed levels of the true differential expression. We recognize the resulting diagnostic as a form of statistical power facilitating interpretation and communication. As a sampling model for observed gene expression densities across genes and arrays, we use a variation of a hierarchical gamma/gamma model. But the discussion of the decision problem is independent of the chosen probability model. The approach is valid for any model that includes positive prior probabilities for the null hypotheses in the multiple comparisons and that allows for efficient marginal and posterior simulation, possibly by dependent Markov chain Monte Carlo simulation.