Power and sample size estimation in high dimensional biology

Power and sample size estimation in high dimensional biology
复制标题

DOI:
10.1191/0962280204sm369ra
复制
发表时间:
2004-01-01
影响因子:
2.3
通讯作者:
Allison, DB
Allison, DB
中科院分区:
医学3区
文献类型:
--
作者:
Gadbury, GL;Page, GP;Allison, DB

文献摘要

被引文献

相似文献

基因组科学家经常在一次实验中测试数千种假设。一个例子是微阵列实验,旨在确定实验组之间的差异基因表达。计划这样的实验需要确定样本量,以便有意义的解释。传统的功率分析方法可能不太适合在以发现为导向的基础研究中测试数千个假设的任务。我们引入了期望发现率(EDR)的概念,以及一种将参数混合建模与参数自举相结合的方法,以估计所需的样本量以达到期望的结果精度。虽然所包括的例子来源于微阵列研究,但这里的方法是研究设计方法中的“超范式”,适用于大多数高维生物情况。来自三个不同微阵列实验的先导数据被用来推断不同样本量和阈值下的EDR以及相关的错误发现率。
Genomic scientists often test thousands of hypotheses in a single experiment. One example is a microarray experiment that seeks to determine differential gene expression among experimental groups. Planning such experiments involves a determination of sample size that will allow meaningful interpretations. Traditional power analysis methods may not be well suited to this task when thousands of hypotheses are tested in a discovery oriented basic research. We introduce the concept of expected discovery rate (EDR) and an approach that combines parametric mixture modelling with parametric bootstrapping to estimate the sample size needed for a desired accuracy of results. While the examples included are derived from microarray studies, the methods, herein, are 'extraparadigmatic' in the approach to study design and are applicable to most high dimensional biological situations. Pilot data from three different microarray experiments are used to extrapolate EDR as well as the related false discovery rate at different sample sizes and thresholds.