Estimating Effect Sizes of Differentially Expressed Genes for Power and Sample-Size Assessments in Microarray Experiments

Estimating Effect Sizes of Differentially Expressed Genes for Power and Sample-Size Assessments in Microarray Experiments
复制标题

DOI:
10.1111/j.1541-0420.2011.01618.x
复制
发表时间:
2011-12-01
期刊:
影响因子:
1.9
通讯作者:
Noma, Hisashi
Noma, Hisashi
中科院分区:
数学3区
文献类型:
--
作者:
Matsui, Shigeyuki;Noma, Hisashi

文献摘要

被引文献

相似文献

在微阵列筛选差异表达的基因,使用多重测试,功率或样本大小的评估是特别重要的,以确保少数相关基因被删除,从进一步考虑过早。在这项评估中,充分估计差异表达基因的效应量是至关重要的,因为它对功效和样本量估计有重大影响。然而,由于随机变异,使用具有最大观察效应大小的顶级基因的常规方法将受到高估。在这篇文章中,我们提出了一个简单的估计方法的基础上分层混合模型与非参数先验分布,以适应随机变化和可能的差异基因,从滋扰,非差异基因的影响大小的大的多样性。基于经验贝叶斯效应量估计,可以估计功效和错误发现率(FDR),以在基因筛选中同时监测它们。我们还提出了一个功率指数,涉及选择最大的效应大小,称为部分功率的顶级基因。这个新的功率指数可以提供一个实际的妥协,在实现高水平的通常的总体功率面临许多微阵列实验的困难。应用程序从癌症临床研究的两个真实的数据集。
In microarray screening for differentially expressed genes using multiple testing, assessment of power or sample size is of particular importance to ensure that few relevant genes are removed from further consideration prematurely. In this assessment, adequate estimation of the effect sizes of differentially expressed genes is crucial because of its substantial impact on power and sample-size estimates. However, conventional methods using top genes with largest observed effect sizes would be subject to overestimation due to random variation. In this article, we propose a simple estimation method based on hierarchical mixture models with a nonparametric prior distribution to accommodate random variation and possible large diversity of effect sizes across differential genes, separated from nuisance, nondifferential genes. Based on empirical Bayes estimates of effect sizes, the power and false discovery rate (FDR) can be estimated to monitor them simultaneously in gene screening. We also propose a power index that concerns selection of top genes with largest effect sizes, called partial power. This new power index could provide a practical compromise for the difficulty in achieving high levels of usual overall power as confronted in many microarray experiments. Applications to two real datasets from cancer clinical studies are provided.